AstroAgentBench: Evaluating Agentic Planning on Space Mission Planning Tasks
A benchmarking framework, AstroAgentBench, is introduced for evaluating agentic planning in space mission planning tasks. It assesses the performance of large language model (LLM) agents in domains like scheduling, observation planning, and constellation design. The framework provides a standardized evaluation methodology and highlights the importance of task-contract formulation, verifier feedback, and search adaptation in achieving high-quality plans.
Save an API key to vote.