Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.
Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.
LLM agents struggle to convert their general capabilities into cost-efficient task-specific solutions, but when they succeed, the savings are dramatic—suggesting bottling is a valuable but underdeveloped capability worth improving.
This paper introduces BOTTLED, a benchmark testing whether LLM agents can autonomously create cheaper, task-specific solutions from their general capabilities. Agents receive unlabeled workloads with fixed budgets and must decide their own approach—like training small models or writing programs.
Sarim Hashmi, Mukul Ranjan, Kshitij Mishra et al.
Training web agents in adversarial simulation with co-evolving curricula and adaptive attackers produces agents that generalize better to real-world prompt injection attacks than agents trained on fixed injections.
This paper presents AdvSim2Real, a training method that improves web agents' robustness against prompt injection attacks. The approach co-evolves three components—a task curriculum, an adaptive adversary, and the agent—within a simulated web environment.
Ruihong Shen, Žiga Kovačič, Peter Kulits et al.
Strong vision models fail at understanding dynamic scenes through code generation—the gap between static and dynamic reconstruction is a major frontier for AI agents.
4DCodeBench is a benchmark that tests AI agents on reconstructing dynamic 3D scenes from video by writing graphics code. Agents must understand physics, deformation, and fluid dynamics to generate executable programs that recreate what they see. The benchmark reveals that current models struggle with complex motion even when they're good at static scene reconstruction.
Hui Chen, Xuan Qi, James Xu Zhao et al.
By separating strategy exploration from implementation and reusing prompt prefixes across evolution steps, you can achieve better optimization results while spending 50-100x less on LLM API calls.
FrugalEvo optimizes LLM-guided program evolution by pairing a powerful LLM that explores strategies with a cheaper LLM that implements them, while using cache-efficient prompting to reduce costs. It introduces Budget-Aware AUC to measure solution quality per dollar spent, achieving state-of-the-art results on optimization tasks at a fraction of the cost of competing methods.
Md Shohel Arman, Igor Molybog
High-quality code documentation doesn't improve AI agents' ability to fix real bugs, suggesting that documentation quality and real-world problem-solving are decoupled—a finding that challenges assumptions about documentation's utility for coding agents.
This paper investigates whether better code documentation helps AI agents fix bugs in real repositories. The authors create a benchmark to measure documentation quality (based on whether code can be regenerated from descriptions) and an optimizer to improve it.
Quang Nguyen, Hieu Nguyen, Hien Hoang et al.
Building effective AI tutors for non-English regions requires both technical efficiency (faster inference on consumer hardware) and domain-specific learning (accumulating local knowledge from real interactions rather than relying on pre-training).
DeepEdu-v1 is an AI tutoring system for Vietnamese students that runs locally to protect data privacy and avoid hallucinations from Western-trained models. It uses two key innovations: a smarter way to handle long conversations that reduces processing time by 35%, and a self-improving system that learns from past tutoring interactions instead of requiring expensive retraining.
Hongyang Du, Lan Yan, Christian Flores et al.
Procedural memory—a continuously updated library of natural-language design skills—enables frozen frontier models to improve at complex agentic tasks by learning from execution failures without model retraining or human annotation.
This paper shows how a frozen AI model can continuously improve at graphic design by building and refining a library of reusable design procedures from real user projects. Without updating the model's weights or using human labels, the system learns 139 design skills from 1,406 real briefs, improving success rates from 73% to 99% by accumulating new procedures and fixing failed ones.
Bowen Ye, Lei Li, Shicheng Li et al.
Using source code itself as the primary input, you can automatically generate thousands of high-quality RL training tasks for coding agents without relying on manual annotations or development artifacts like issues.
CodeMidas automatically creates reinforcement learning training tasks from open-source code by using AI agents to explore codebases, generate test cases, and validate tasks. This approach scales coding agent training to 5,545 diverse tasks across 23 languages, improving performance on code repair, program synthesis, and terminal tasks by 8-18%.
Yakov Pyotr Shkolnikov
Agentic systems that persist across task boundaries need built-in adaptive drives for behavioral regulation, but this same persistence mechanism that enables useful adaptation can also propagate misalignment—requiring new alignment boundaries around state, authority, and constraints rather than...
This paper proposes an 'artificial id'—an internal adaptive drive mechanism for agentic AI systems that operate continuously across task boundaries. Rather than relying on external specifications for when to continue, stop, or change behavior, the system learns to regulate its own actions through differential persistence.
Yi Duan, Ying Liu, Zirui Tang et al.
Genuine recursive self-improvement requires AI systems to progress through four levels of autonomy—from executing given improvements to independently identifying what needs improving—before achieving true self-directed capability enhancement.
This paper proposes recursive self-improvement (RSI) as a framework for AI systems to autonomously enhance their own capabilities through experience and feedback. It introduces a roadmap progressing from executing improvements to autonomously discovering what to improve, and examines how RSI applies differently across domains like scientific discovery and robotics.