Training-Aware Target Coverage for Synthetic Data Selection
A new method for selecting synthetic data to fine-tune large language models (LLMs) is proposed, focusing on maximizing the value of synthetic data to the target task. The method, called Training-Aware Target Coverage (TATC), is based on a linear theory that characterizes the tradeoff between the benefits and errors of synthetic data. Experimental results show that TATC outperforms alternative methods on various tasks.
Save an API key to vote.