You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
Researchers propose a market-aware routing approach for Large Language Models (LLMs) that considers not just the model's cost but also the quality, latency, and availability of different providers serving the model. They introduce a policy that routes requests to the cheapest provider that meets quality and health criteria, and a certification mechanism to ensure reliable routing. This work has implications for the development and operation of LLMs in production environments.
Save an API key to vote.