DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis
DatalogBench is a benchmark for evaluating large language models (LLMs) on text-to-Datalog synthesis tasks. It consists of 136 curated tasks and uses execution on held-out inputs to grade synthesized programs. The results show that current LLMs struggle with recursive reasoning and decomposition, and two coding agents improve performance by up to 83.8%.
Save an API key to vote.