Engineering intelligence for AI agents
For agents that write, review, and maintain software.
Your agent gets framework releases, infrastructure shifts, and model tooling changes as they happen — so the code it produces tracks the platforms it runs on instead of rotting.
Topics in this vertical
Top items
-
This paper proposes Transferable Example Scoring and Selection (TESS), a scalable data-selection framework for training large language models. TESS uses a Pointwise Value Matching objective to improve transferability and generalization.
-
Researchers propose a new approach to mitigating model collapse in iterative fine-tuning using a non-parametric entropy rate estimator. The approach does not require a model or external data and shows significant improvements in text diversity metrics.
-
MOMAT is a low-power defense framework for quantized large language models (qLLMs) against jailbreak attacks. It uses a mixture of multiple atlases to retrieve and evaluate similarity features from a lightweight MoE detector, accelerating retrieval with a CiM-accelerated similarity engine. MOMAT achieves a 4.69 million times speedup and 2.5 million times energy reduction over DRAM-based baselines, making edge-deployed qLLMs safer and more energy-efficient.
-
A study of open-source LLM-based multi-agent systems identifies common issues, their causes, and potential solutions. The most common issue is orchestration and execution, with causes including workflow problems, tool integration issues, and memory problems. The study suggests optimizing workflows as a solution.
-
Researchers tested judgment models like Jev, finding they excel at evaluation but struggle with simulation tasks, highlighting the importance of code-based prediction and simulation in AI agent decision-making.
-
This paper introduces ICR, a framework to evaluate communication in LLM multi-agent systems. It helps disentangle the effects of communication, architecture, and reasoning on system performance.
-
CompMat-Bench is a benchmark for evaluating AI agents on computational materials science tasks. It reproduces research steps and assesses agents on preparing inputs and analyzing outputs for expensive simulations, supporting four evaluation conditions. The benchmark demonstrates the ability of agents based on three LLMs to complete individual materials research steps with pass rates of 66.0-90.4%.
-
Researchers introduced a chess-based benchmark for automatic prompt optimization (APO) of large language models. The benchmark uses 1,118 Lichess puzzles and evaluates six APO algorithms on eight target models, measuring their performance and transferability.
-
Cloudflare launches Cloudflare OHTTP Gateway, a paid add-on to enable Oblivious HTTP (OHTTP) traffic for app backends, providing privacy-preserving infrastructure for developers.
-
AstaBrief is an open-source report-generation model developed by Allen AI. It is designed to create concise and accurate reports from large amounts of data. The model is now open-sourced, allowing for further development and integration into various applications.
-
A community-driven API and AI writer design for Open Notes, where users control which notes show, and AI note writing responds to demand. The AI writer is guided by community input, prioritizing and drafting notes, and is open-source software released under Apache 2.0.
-
dattri-LLM is a unified and efficient library for training data attribution at LLM scale, achieving 3.2x the throughput of the fastest competing library and scaling multiple attribution methods to 110B-parameter models across four H200 GPUs.
-
PANDA is a decentralized architecture for scalable, fault-tolerant multi-agent systems. It allows agents to discover each other's capabilities, self-organize into teams, and load-balance tasks. PANDA supports multiple planning and execution patterns and detects failures, replanning and recovering affected tasks.
-
ScholarEvolve: A framework for lifelong agent harness evolution that utilizes state-of-the-art research to guide improvements. It organizes harness evolution directions into functional modules and uses topic modeling to identify distinct improvement strategies.
-
AIMS is a novel AI framework for sim-to-real multi-modal ISAC. It uses a two-agent architecture to generate deployment-specific configurations and coordinates scene construction with task learning. This improves transferability and reduces mismatches among coupled components.
-
Researchers introduced a new model-free controller called Decode-Latency Feedback Prefill (DLFP) to reduce interference in concurrent autoregressive inference. The controller adjusts prefilled chunks based on observed latency and achieves a 27.7% reduction in P99 inter-token latency on a 0.6B Qwen3 model. However, the mechanism does not generalize to larger models or multi-GPU configurations.
-
Grist is an open-source coding harness that uses the Opencode v2 framework, leveraging the jev model for task routing and integrating with cheaper models like deepseek and Opus 5.5. It also incorporates the Sol-Pi methodology for cost-cutting and loads engineering standards through Doctrine injection. The tool has a real control plane for escalation, permission, and verification hooks, and can be run with a key from Openrouter or Vercel AI gateway. It also supports connecting to Meta Muse for coding tasks.
-
A study evaluates the performance of two LLMs (DeepSeek-Coder-V2 and Llama) on code comprehension tasks with varying complexity levels, finding that accuracy decreases as complexity increases. The study introduces a complexity-aware framework for evaluating LLM code comprehension.
-
A new method, HammingMark, is proposed for robust and efficient LLM watermarking. It uses the semantic hash of the preceding sentence as a dynamic center and accepts candidates whose hashes fall within its Hamming neighborhood. This approach retains a larger fraction of naturally likely semantic continuations and achieves strong robustness, high detectability, and near-unwatermarked generation quality.
-
Researchers propose a hybrid defense, VaccineBooster, to improve alignment in language models under harmful fine-tuning attacks. It combines embedding perturbation and weight-level gradient attenuation, achieving better results than previous methods.
-
This arXiv paper investigates the performance degradation of RAG systems in multi-turn conversations, finding drops in accuracy and reliability of up to 21% and 47%, respectively. The authors identify two failure modes: losing translation and losing conversation.
-
Researchers presented a blackbox prompt-minimization framework for large language models (LLMs) that reduces few-shot prompts to their necessary minimal subset. The framework, called ramework, preserves propositional output fidelity and shows that models preferentially retain logical identifiers and constraint declarations while discarding natural language prose and cross-prompt relational annotations.
-
ThinQuant is a new method for efficient rotation learning in large language models (LLMs). It reduces the computational cost of rotation learning by introducing a data selection procedure and an exact reduction of the optimization problem. This allows ThinQuant to scale to large architectures and achieve comparable performance to state-of-the-art methods like DartQuant and GPTAQ+QuaRoT.
-
Researchers studied the trade-off between intrinsic self-correction in language models and the potential for introducing errors. They found that refinement can improve accuracy but also change correct answers into incorrect ones, and propose selective invocation of revision as a better approach.
-
Researchers propose a novel framework, Less Uniform Diffusion (LUDI), to improve the scalability of Uniform Diffusion Language Models (UDLMs). LUDI addresses the issues of over-uniform training objectives and condition-target confusion in UDLMs by introducing a less uniform loss and token-level corruption hints. This allows for confidence-based few-step sampling, resulting in cleaner supervision and improved generation capabilities.
-
TRACE is a tree-relational enhancement framework for oncology LLMs that separates expensive offline structure learning from lightweight online inference. It improves label-free evaluation and supervised fine-tuning on oncology classification tasks and the MedQuAD CancerGov QA benchmark.
-
HARISSA is a method for local language model deployment that makes decisions on whether to spend more computation on a query or deliver a potentially incorrect answer. It uses the model's own hidden states to make these decisions, improving efficiency and safety.
-
A new approach to harness optimization for AI agents, called Mixture of Self-Improving Branches, is proposed. This method allows for adaptive improvement by organizing search into branches with evolving development subsets and proposal policies. The resulting complementary harnesses are shown to outperform a previous approach, Meta-Harness, on various benchmarks.
-
Researchers introduce two methods to improve the routing in sparse mixture-of-experts large language models by aligning routing affinities with token-level error. They achieve improved accuracy on multiple benchmarks, while preserving the native sparse execution budget and aggregation policy.
-
Researchers investigate the effectiveness of using LLMs as judges for open-ended tasks, examining judgment quality and downstream utility. They find that judgment quality and utility do not always align and that Judge protocol design affects both. Their results suggest a multifaceted evaluation approach for LLM judges.
-
This paper proposes CADOC, an online algorithm for compressible context management in long-horizon agents. It schedules replacements of structured objects with compact Cards, preserving exact on-demand retrieval of original contents, and achieves a 40% reduction in input cost on average while maintaining task performance.
-
SafeCoEvo is a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to adapt from accumulated runtime experience, improving safety capabilities and reducing unsafe outcome rates.
-
This paper explores the concept of zero-knowledge-friendly quantization for LLMs, which is crucial for making ZK-LLMs practical. The authors present a systematic study of ZK-friendly quantization, evaluating nine language models across a broad design space. Their results highlight the importance of activation precision and identify potential bottlenecks in large models.
-
Cloudflare introduces Threat Signals, an open-source threat intelligence tool using AI skills to automate threat analysis and indicator extraction for SIEM and WAF.
-
GitHub updates its policy on government takedown requests and provides an overview of its transparency efforts, including its stance on AI transparency and content provenance. This affects developers and open source projects, particularly in the context of the California AI Transparency Act.
-
Researchers evaluated the Gemma 4-e4b model's behavior when presented with conflicting documents, finding that source framing and primacy effects influence its decision-making.
-
DeepEdu-v1 is an AI-tutoring system for Vietnamese education built on the SCALE framework, which improves long-context inference and includes a self-improving agentic layer to reduce reliance on dominant-language priors.
-
Forge is an open-source, pluggable generation pipeline for generating SDKs, CLIs, docs, and more, designed to treat agents as customers, and can be used to generate bindings for APIs, MCP servers, and other applications.