Benchmarking Prompt Optimization of Large Language Models With Chess

Researchers introduced a chess-based benchmark for automatic prompt optimization (APO) of large language models. The benchmark uses 1,118 Lichess puzzles and evaluates six APO algorithms on eight target models, measuring their performance and transferability.

RSS Score 0 10/2/2026, 4:00:00 AM Original Source
Save an API key to vote.