Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy
A research paper evaluates the ability of large language models (LLMs) to reconstruct implicit scientific knowledge in astronomy by reproducing published research results. The paper proposes a framework for end-to-end reproduction, separating execution from verification, and highlights the limitations of current LLM-based agents in recognizing causal relationships in implicit knowledge.
Save an API key to vote.