HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment
HARISSA is a method for local language model deployment that makes decisions on whether to spend more computation on a query or deliver a potentially incorrect answer. It uses the model's own hidden states to make these decisions, improving efficiency and safety.
Save an API key to vote.