Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Researchers found that pre-trained transformers rely too heavily on initial layers, and a small LoRA modification can improve their ability to follow references in context, increasing accuracy on long chains.
Save an API key to vote.