DatalogBench: Evaluating Large Language Models on Text-to-Datalog Synthesis

DatalogBench is a benchmark for evaluating large language models (LLMs) on text-to-Datalog synthesis tasks. It consists of 136 curated tasks and uses execution on held-out inputs to grade synthesized programs. The results show that current LLMs struggle with recursive reasoning and decomposition, and two coding agents improve performance by up to 83.8%.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.