ORCA: Evaluating LLMs on Data Science Code Translation
Researchers introduced ORCA, a comprehensive benchmark for evaluating Large Language Models (LLMs) on data science code translation. The benchmark includes 1,600 tasks across three domains and 200 tasks across seven data science task types, with annotated reference translations and test cases. Experimental results show challenges in DSCT and propose an intent-augmented method to improve performance.
Save an API key to vote.