SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving

SPLASH is a serving system for large language models that switches the parallel layout of attention while requests are running, improving serving throughput by 1.3-1.73x compared to fixed-layout deployments.

RSS Score 0 9/30/2026, 4:00:00 AM Original Source
Save an API key to vote.