ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2612 papers42 this month12 topics
AllTraining 50Efficiency 38Reasoning 37Agents 32Evaluation 32Architecture 18Applications 13Safety 12Multimodal 11scaling 6Data 5Alignment 3

Oct 5 – Oct 11(17)

BrickBench: Evaluating Agentic Brick Design

Oct 8, 2026

Peter Kulits, Yiqing Xu, R. Kenny Jones et al.

Current leading AI agents can meet basic physical and semantic requirements for LEGO design but struggle to match human-level design quality, revealing gaps in spatial reasoning and constraint satisfaction under discrete choice problems.

BrickBench is a benchmark that tests AI agents' ability to design buildable LEGO sets from text descriptions. Agents must select parts from a library and satisfy physical constraints, semantic requirements, and design quality—combining reasoning about local details with global structure. The paper includes BrickAgent, an environment for agents to build and validate designs.

agentsreasoningevaluation

RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments

Oct 8, 2026

Zimo Wen, Yijin Chen, Yuxuan Cao et al.

By structuring robot learning around explicit skill hierarchies with clear input-output contracts and execution-grounded diagnosis, robots can reliably improve their capabilities through experience and safely reuse learned skills across new tasks.

RoboRSI is a robot self-improvement system that learns and refines skills through real-world experience. It organizes task execution into a hierarchy of skills (compound, atomic, base) with clear responsibilities, diagnoses failures to pinpoint which skill needs fixing, and validates improvements before reusing them.

Sep 28 – Oct 4(45)

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Oct 2, 2026

Ruihong Shen, Žiga Kovačič, Peter Kulits et al.

Strong vision models fail at understanding dynamic scenes through code generation—the gap between static and dynamic reconstruction is a major frontier for AI agents.

4DCodeBench is a benchmark that tests AI agents on reconstructing dynamic 3D scenes from video by writing graphics code. Agents must understand physics, deformation, and fluid dynamics to generate executable programs that recreate what they see. The benchmark reveals that current models struggle with complex motion even when they're good at static scene reconstruction.

evaluationreasoningagents

What Should World Models Forget? Stratified Retention for Continual Adaptation

Oct 2, 2026

Nishit Anand, Ramani Duraiswami, Dinesh Manocha

World models need stratified forgetting strategies that preserve physical invariants while quickly adapting to environmental changes—standard continual learning metrics fail to capture this distinction and incorrectly reward frozen models.

This paper addresses a fundamental problem in continual learning for world models: knowing what to forget. Unlike traditional learning where correct labels stay correct, world models operate in changing environments where outdated knowledge must be discarded.

Sep 21 – Sep 27(32)

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

Sep 25, 2026

Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe et al.

Training language models to predict their confidence in intermediate reasoning steps—using only self-supervised learning—makes them generate shorter reasoning traces at inference time without any explicit length penalties or early-stopping mechanisms.

This paper shows that reasoning models can generate shorter, more efficient reasoning traces by learning to predict their own confidence in answers—without explicitly optimizing for length.

trainingefficiencyreasoning

First-Order Stationarity of Reverse Diffusions

Sep 25, 2026

Zhifeng Chen, Chenyang Jiang, Yazhen Wang

Diffusion models have provable convergence guarantees similar to optimization algorithms—reverse diffusions contract divergence exponentially fast, and discrete samplers achieve measurable stationarity bounds that don't depend on data convexity.

This paper connects optimization theory to diffusion models by proving that reverse-time diffusion processes contract Fisher divergence at exponential rates under strong convexity conditions. The authors also establish first-order stationarity bounds for practical discrete samplers, showing how optimization guarantees translate to sampling quality without requiring global convexity.

Sep 14 – Sep 20(6)

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Sep 18, 2026

Bowen Ye, Lei Li, Shicheng Li et al.

Using source code itself as the primary input, you can automatically generate thousands of high-quality RL training tasks for coding agents without relying on manual annotations or development artifacts like issues.

CodeMidas automatically creates reinforcement learning training tasks from open-source code by using AI agents to explore codebases, generate test cases, and validate tasks. This approach scales coding agent training to 5,545 diverse tasks across 23 languages, improving performance on code repair, program synthesis, and terminal tasks by 8-18%.

trainingagentsreasoning

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

Sep 18, 2026

Andre Bacellar

Multi-hop retrieval failures follow predictable structural patterns that vary by dataset; you can detect high-risk queries using simple features (query length, retrieval concentration) and a learned confidence score, enabling safe abstention without additional LLM calls.

This paper identifies why multi-hop retrieval systems fail predictably on certain queries and proposes a method to detect these failures without extra LLM calls. The authors prove that failure patterns cluster in specific subpopulations and that different query features predict failure in different scenarios.

agentsreasoningtraining

Decoupling Exploration from Optimization in RLVR

Oct 7, 2026

Saif Punjwani, Micah Goldblum

Decoupling exploration (with novelty bonuses) from optimization (standard training) via distillation lets language models discover diverse correct reasoning strategies without quality degradation, outperforming direct RLVR approaches.

This paper proposes Exploration-Distillation (ExpDis), a method that separates exploration from optimization in reinforcement learning with verifiable rewards. Explorer policies use novelty bonuses to discover new reasoning strategies, their best trajectories are filtered and distilled into a student policy trained without novelty incentives.

trainingreasoningagents

RoboJEPA: Scaling Robotic Latent World Models

Oct 7, 2026

Artem Zholus, Nicolas Beltran-Velez, Jianhao Yuan et al.

Latent world models follow reliable scaling laws: imagination error (how well the model predicts future states) improves predictably with compute, and this directly correlates with real robot task performance, making it possible to forecast model quality without expensive robot experiments.

RoboJEPA is an 8-billion-parameter world model trained on real robot data from 12 different embodiments. It predicts future states in a compressed latent space and follows predictable scaling laws—meaning you can estimate how much better it gets with more compute before actually training it. The model works as a zero-shot robot controller, planning by imagining future states toward a goal image.

scalingreasoning

SciExam for ENSO: Can AI Agents Build Climate Models?

Oct 7, 2026

Yinling Zhang, Langchen Liu, Dongbin Xiu et al.

AI agents can conduct genuine scientific research on unsolved problems—this benchmark shows agents building climate models that outperform published work and align with competing scientific theories, proving evaluation is possible without a known ground truth.

Researchers created SciExam for ENSO, a benchmark where AI agents build climate models of El Niño from real data without knowing the right answer. Agents work within a 6-hour budget, write their own diagnostics, and develop models tested against hidden criteria.

agentsreasoningevaluation

RECAST: Learning to Compute the Right Context through Adaptive Evidence Routing

Oct 7, 2026

Yilun Hao, Krishna Sayana, Isabella Ye et al.

Instead of just retrieving relevant documents, RECAST learns to combine retrieval with computation—filtering, aggregating, and deriving answers across multiple sources—enabling LLMs to handle complex, multi-step reasoning tasks more effectively.

RECAST is a framework that helps language models solve complex tasks by learning to actively construct evidence through computation rather than just retrieving it. A lightweight router model decides which operations to perform on multiple information sources, a compiler translates those decisions into executable code, and an answer model produces the final result.

reasoningagents

EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution

Oct 7, 2026

Python Song, Zhixuan Liang, Kelsey Fu et al.

By treating robot learning as active hypothesis testing rather than passive data collection, you can dramatically reduce the physical experiments needed to improve foundation models—this system reaches 77% success where baselines only achieve 40%.

EmbodiedRSI is a self-evolving robot control system that autonomously decides which experiments to run on physical robots and uses the results to improve its code and skills.

agentsreasoningtraining

Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective

Oct 6, 2026

Kevin Zhang, Stephen Bates

Conformal prediction set sizes have a rigorous information-theoretic interpretation: they quantify information gain in a way that's mathematically sandwiched between generalized entropy measures and obeys data processing inequalities.

This paper establishes a theoretical connection between conformal prediction (a method for uncertainty quantification) and information theory. The authors show that the size of prediction sets from conformal methods can be interpreted as a measure of information gain, providing mathematical justification for using set size as an uncertainty metric.

evaluationsafetyreasoning

4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

Oct 6, 2026

Shiqi Li, Sean Cho, Yijie Li et al.

Feed-forward generative models can reconstruct complex hand-object interactions faster and more reliably than optimization-based methods by learning to correct foundation model errors while respecting physical constraints during generation.

This paper presents 4D-HOF, a fast method for reconstructing 3D hand and object positions/orientations from video. Instead of slow per-video optimization, it uses a generative model trained on diverse data to refine rough estimates from vision models.

architectureefficiencyreasoning

IdeaAnchor: Teaching LLMs to Turn Literature into Research Ideas

Oct 6, 2026

Ziyu Chen, Yilun Zhao, Jiashuo Sun et al.

Structured supervision—encoding how papers should be synthesized together—is more effective for training ideation models than prompting alone, and combining anchor-based training with retrieval produces the best results.

This paper teaches language models to generate research ideas by synthesizing multiple papers, using structured specifications called 'anchors' that capture how papers should be combined.

trainingreasoningapplications

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

Oct 6, 2026

Zewei Zhou, Rachel Luo, Yulong Cao et al.

Fixing your evaluator is as important as fixing your policy—when agents improve, what they need to be judged on changes, so you need a system to evolve your judge alongside your agent.

VeriFine is a framework that improves AI agents through co-evolving the policy, training data, and evaluation judge. When an agent's performance plateaus, humans help refine the judge by resolving disagreements on tricky cases, then the improved judge guides better training. Tested on driving and robot navigation, it shows continuous improvement as new failure patterns emerge.

trainingreasoningagents

Linear Bandits under Exact Sliding-Window Constraints

Oct 6, 2026

Seyed Mohammad Hadi Hosseini, Yasin Abbasi-Yadkori, Sattar Vakili

Sliding-window constraints make online learning fundamentally harder than offline optimization, but rare policy updates combined with optimistic planning can achieve sublinear regret while maintaining exact feasibility.

This paper studies how to make optimal decisions in linear bandit problems when actions must satisfy strict sliding-window constraints—meaning every consecutive block of actions must come from an allowed set.

trainingreasoningefficiency

On the Computational Tractability of Robust Bandits

Oct 6, 2026

Vanessa Kosoy, Vinayak Pathak

Robust bandits have a sharp computational boundary: a specific case is tractable with efficient algorithms, but small generalizations become NP-hard, suggesting this marks the frontier of what's computationally feasible for unrealizable learning.

This paper studies how to efficiently learn in bandit problems when the true environment doesn't match the learner's model. The authors identify a tractable special case with polynomial-time algorithms and $\tilde{O}(\sqrt{T})$ regret, while showing that natural generalizations become NP-hard—establishing where the problem transitions from solvable to hard.

trainingreasoningsafety

Base Models Can Reason By Taking a Cue From Training Data

Oct 5, 2026

Sophie L. Wang, Amil Dravid, Rulin Shao et al.

Base models already contain reasoning capabilities encoded in their training data—you can unlock them by conditioning on the right token cues, without needing expensive RL fine-tuning.

This paper shows that base language models can achieve reasoning performance comparable to RL-trained models by using specific starting tokens (like "Okay" or "Alright") that trigger learned associations from training data. The authors demonstrate they can create new reasoning cues through data interventions and trace these effects back to specific document types in the training set.

trainingreasoningdata

Recursive Video In-Context Learning for Agentic Robot

Oct 5, 2026

Wenrui Bao, Xinxin Liu, Bingxin Xu et al.

By organizing demonstration videos into navigable hierarchies rather than static prompts, agents can access task details only when needed, reducing context overhead while improving learning from single examples.

This paper presents Recursive Video In-Context Learning (RV-ICL), a method that helps robot agents learn from demonstration videos more effectively. Instead of feeding entire videos as prompts, RV-ICL organizes a single demo video into a hierarchical structure of sub-events (like grasps and releases) that the agent can navigate on-demand.

agentsreasoningmultimodal

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Oct 5, 2026

Yifan Zhang, Yutong Dai, Viraj Prabhu et al.

Self-verification through conformal methods lets web agents learn from their own reasoning about task progress, eliminating the need for expensive judge calls at deployment while improving training efficiency.

CLIFT trains web agents to complete browser tasks by having them verify their own actions through natural-language questions, creating a reusable signal that works both during training (with sparse rewards) and at test time (without expensive external judges). The method achieves state-of-the-art results on multiple web agent benchmarks and transfers across different models.

trainingagentsreasoning

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

Oct 5, 2026

Oliver Jaffe, Dane Sherburn

AI models are becoming exponentially better at the experimental research process itself (not just raw capability), with frontier models now reaching expert-level results using 2.3x less compute—a skill that could significantly accelerate AI R&D timelines.

TasteVal is a benchmark measuring how efficiently AI models design and conduct experiments to solve research problems. Rather than evaluating raw problem-solving ability, it measures 'experimental taste'—the skill to iteratively design good experiments and interpret results—by comparing how much compute a model needs versus human experts to reach the same performance level.

evaluationreasoningagents
trainingevaluationreasoning

Language Models that Play Chess and Explain Their Moves

Oct 2, 2026

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Language models can match expert-level performance in specialized domains by distilling knowledge from silent expert systems through iterative refinement, opening a path to explainable AI in games, robotics, and other domains with strong baseline models.

This paper presents Queen, a 4-billion-parameter chess model that combines a silent chess engine with a language model to play at Grandmaster level while explaining its moves.

reasoningtrainingapplications

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

Oct 2, 2026

Hui Chen, Xuan Qi, James Xu Zhao et al.

By separating strategy exploration from implementation and reusing prompt prefixes across evolution steps, you can achieve better optimization results while spending 50-100x less on LLM API calls.

FrugalEvo optimizes LLM-guided program evolution by pairing a powerful LLM that explores strategies with a cheaper LLM that implements them, while using cache-efficient prompting to reduce costs. It introduces Budget-Aware AUC to measure solution quality per dollar spent, achieving state-of-the-art results on optimization tasks at a fraction of the cost of competing methods.

efficiencyreasoningagents

Planning to Learn

Oct 2, 2026

Ian Osband

Cross-entropy beats policy gradients in classification because it's 'patient'—it optimizes for total error reduction across all future steps, not just immediate accuracy. A simple horizon-aware loss can capture this benefit while staying closer to principled gradient methods.

This paper reveals why cross-entropy outperforms exact policy gradients in classification despite having access to the true label. The key insight is that cross-entropy implicitly accounts for future learning steps, while exact policy gradients are myopic.

trainingreasoning

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Oct 2, 2026

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi et al.

By identifying and selectively training on high-impact token decisions rather than full sequences, you can make diffusion language models learn more efficiently with less data.

This paper introduces Pivot-SD, a training method for masked diffusion language models that focuses on the most impactful decisions during text generation. Instead of training on entire sequences, it identifies 'pivot' tokens—commitments that significantly reduce uncertainty about remaining words—and trains only on those, using success/failure signals to guide learning.

trainingefficiencyreasoning

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Oct 2, 2026

Vladislav Gromadskii, David Li, Samson Gourevitch et al.

You can fine-tune fast diffusion generators for reward optimization without expensive reference rollouts by replacing intractable KL penalties with inverse-distillation regularization that provably bounds divergence.

IDRF is a method for fine-tuning masked discrete diffusion models (which generate sequences iteratively by predicting multiple tokens at once) to maximize rewards while staying close to a reference model. Instead of computing intractable likelihood penalties, it uses a clever regularization trick called inverse-distillation that upper-bounds the KL divergence.

trainingefficiencyreasoning

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Oct 1, 2026

Yen-Jen Wang, Haozhe Jiang, Shuying Deng et al.

Robots can improve their own performance through autonomous practice and skill refinement in simulation without updating model weights, then transfer successfully to real hardware—a practical path to reliable robot systems.

RPG is a framework that improves robot performance without retraining models by identifying skills from offline data, practicing in simulation with failure diagnosis, and refining symbolic skills and system prompts.

agentsreasoningtraining

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Oct 1, 2026

Sohyeon Kim, Yoonho Lee, Bo Liu et al.

Even advanced AI agents fail at retrieving papers that inspired real research (max 0.51 recall), revealing a critical gap in how models search scientific literature—this task requires something beyond current retrieval and reasoning approaches.

ScholarCatalyst is a benchmark dataset where 184 computer science researchers labeled which prior papers inspired their completed projects. The benchmark tests whether AI systems can retrieve these influential papers given only an initial research question and literature available at project start.

evaluationreasoningagents

VISTA: A Visual Harness for Reasoning in an Interactive World

Oct 1, 2026

Qiushi Han, Keya Hu, Linlu Qiu et al.

Adding a simple visual memory system that preserves and lets models retrieve past observations dramatically improves multimodal models' reasoning in interactive visual environments—achieving perfect performance on challenging puzzle games.

VISTA is a visual harness that enhances multimodal models' ability to solve complex interactive visual tasks by giving them long-horizon vision and lossless visual memory. The system lets models directly perceive environments, store past observations, and actively retrieve them while reasoning.

agentsmultimodalreasoning

FERPO: Forward Entropy-Regularized Policy Optimization

Oct 1, 2026

Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

Using forward-KL instead of reverse-KL for policy fitting encourages broader exploration of high-value actions and avoids the computational cost of differentiating critics, leading to faster and more sample-efficient learning.

FERPO is a reinforcement learning algorithm that improves policies by using critic values directly rather than differentiating through them. It derives optimal target actions using entropy regularization and fits the actor to these targets using forward-KL divergence, which encourages exploring multiple high-value action modes while keeping importance weights stable.

trainingefficiencyreasoning

Hierarchical Continuous Diffusion Language Models

Oct 1, 2026

Hui Ren, Zihan Li, Chang Liu et al.

HC-DLM bridges discrete and continuous diffusion by coupling token generation with a shared latent trajectory, enabling better reasoning and constraint satisfaction than purely discrete or continuous approaches.

This paper proposes Hierarchical Continuous Diffusion Language Models (HC-DLM), which combines discrete token generation with continuous latent states in a single denoising process. Unlike existing approaches that treat these separately, HC-DLM uses the continuous latent as the only persistent state, reading tokens from it at each step.

architecturereasoningtraining

Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control

Oct 1, 2026

Akshay Balsubramani

Cost-augmented optimal transport on graphs can be solved exactly and efficiently using matrix exponentials instead of learned neural networks—no approximation or temporal discretization needed.

This paper solves the cost-augmented Schrödinger bridge problem on graphs exactly, without learning or time discretization. By reformulating state costs as a Feynman-Kac tilt of the reference process, the problem reduces to computing a standard bridge through alternating matrix exponentials. The method is exact, memory-efficient, and converges based on endpoint coupling alone.

reasoning

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Oct 1, 2026

Shuo Xing, Zilin Dai, Chengyuan Qian et al.

Mathematical reasoning in LLMs isn't a single skill but four distinct capabilities; focusing training on the 'Discovery' bottleneck (finding the right solution strategy) is more effective than generic math training.

This paper diagnoses why LLMs struggle with math by breaking down mathematical reasoning into four components (Discovery, Generation, Digestion, Execution) and shows that Discovery—finding the right approach—is the main bottleneck. The authors then propose a training method that uses these insights to improve math performance across different model sizes.

reasoningtrainingevaluation

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Oct 1, 2026

Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.

Sparse autoencoder features combined with reinforcement learning enable interpretable, composable control over protein sequence generation—activating 3.75x more targeted features than previous steering approaches and improving predicted biological function.

IDiom is a specialized protein language model trained on 54 million intrinsically disordered protein regions (IDRs) that can generate functional sequences with precise control over biological features.

trainingapplicationsreasoning

Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination

Oct 1, 2026

Suyu Ye, Zheyuan Zhang, Vaishnav Tadiparthi et al.

Observing joint behavior between two robots reveals hidden physical constraints better than observing a single constrained robot, enabling effective zero-shot coordination without explicit communication.

This paper tackles a practical robotics problem: how can one robot learn another robot's physical limitations (like broken joints or weak actuators) just by watching them work together, then use that knowledge to coordinate effectively on new tasks? The authors show that by analyzing how both robots move together, you can infer hidden constraints better than looking at just one robot alone.

agentsreasoning

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Oct 1, 2026

Xuan Zhang, Longtao Zheng, Cunxiao Du et al.

Teaching agents to manage their own context through learned compaction decisions—rather than just handling overflow—improves performance on long-horizon coding tasks by 5-9% across different context window sizes.

AutoCompact trains coding agents to automatically decide when and how to compress their working context during long software engineering tasks. By learning when to discard stale exploration and what state to preserve, the agent improves its ability to solve repository-level coding problems while staying within context limits.

agentsreasoningtraining

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Oct 1, 2026

Hanchu Zhou, Dechen Gao, Hang Wang et al.

Using semantic communication between robots—where they exchange meaningful descriptions rather than sensor data—enables better coordination on long-horizon tasks while keeping each robot's execution independent and reliable.

DuoMind is a framework that enables multiple robots to coordinate and work together on complex tasks by combining vision-language models for high-level reasoning with vision-language-action models for precise execution. Robots communicate through semantic messages rather than raw data, allowing them to share understanding of the task and environment while maintaining independent control.

agentsmultimodalreasoning

When Do Intrinsic Rewards Lead to Exploration?

Oct 1, 2026

Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett

Intrinsic rewards designed to encourage exploration can fail to find the most informative experiences—you need to explicitly measure whether an agent's history can substitute for real experience under different policies.

This paper examines when intrinsic reward signals (like prediction error or curiosity) actually lead to good exploration in reinforcement learning. The authors show that maximizing these rewards doesn't always produce the most informative experiences, propose a formal criterion based on counterfactual information, and demonstrate failures of existing methods with concrete examples.

trainingreasoning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Oct 1, 2026

Lucheng Fu, Kejing Xia, Yiyang Wang et al.

LLM agents can significantly improve performance on knowledge-intensive tasks by learning persistent, source-specific models that evolve through repeated interaction—achieving up to 22.6 point gains over standard retrieval methods.

This paper introduces SourceLearn, a method for LLM agents to develop persistent, reusable understanding of external knowledge sources through repeated interaction.

trainingagentsreasoning

Faynt: Scaling and Optimizing Policies for Competitive Melee

Oct 1, 2026

Ali Janati, Nikita Kuzmin, Rohit Swamy et al.

Scaling, architecture choices, and post-training curricula matter more than model size alone—a smaller, optimized policy outperforms a larger pretrained one, suggesting careful design beats raw parameter count for complex game-playing tasks.

Faynt introduces transformer-based AI policies for Super Smash Bros. Melee that control all 26 characters with a single model. A 10M-parameter version wins 98.4% of same-character matches against existing AI opponents and beats a zero-delay competitor, using techniques like supervised pretraining on 840k human replays, curriculum learning, and distillation from a larger 75M model.

trainingreasoningagents

Sample complexity bounds for categorical Markov random fields via Discrete Diffusions

Oct 1, 2026

Shivam Kumar, Nabarun Deb

Discrete diffusions can sample from structured categorical data with provable sample complexity guarantees, and weight-sharing neural networks learn scores more efficiently than fully-connected ones for this task.

This paper develops theoretical guarantees for sampling from high-dimensional categorical distributions using discrete diffusion models.

trainingreasoningscaling

GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

Oct 1, 2026

Yakun Zhu, Yi Bin, Yujuan Ding et al.

Separating geometry into common and residual components while routing visual learning through structured latents prevents representation collapse and improves 3D spatial reasoning from images.

This paper addresses 3D spatial reasoning from 2D images by introducing GeoLatent, which uses decomposed spatial representations (position, direction, geometry) with geometric supervision and routed optimization. The method prevents geometry collapse and ensures latents are actively used during learning, achieving state-of-the-art results on spatial reasoning benchmarks.

reasoningmultimodalarchitecture

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

Oct 1, 2026

Arman Behnam, Binghui Wang

Current memory systems can't measure the value of memories that are never retrieved.

Memory-augmented language models struggle to identify which memories are actually useful because some memories are never retrieved, making their value impossible to measure.

reasoningevaluationtraining

AI Emulation of Stochastic Sudden Stratospheric Warming with Interpretable Latent Structure

Oct 1, 2026

C. Daniel Boscu, Daniel Hernandez, Fabio Alvarez Ventura et al.

Deep learning models can be designed to both predict rare, extreme events accurately and produce interpretable internal representations that align with real physical phenomena, enabling better understanding of how AI makes decisions about complex systems.

Researchers built an AI model to predict rare weather events called Sudden Stratospheric Warming by learning from a simplified atmospheric model.

reasoningevaluationapplications

Semifactual Credit-Augmented Policy Optimization

Sep 30, 2026

Junshu Pan, Zhizhang Fu, Shulin Huang et al.

Token-level stability under prompt variations is a useful training signal for improving reasoning in LLMs—you can boost performance by penalizing tokens that change meaning when irrelevant prompt details change.

This paper identifies that language models trained with reinforcement learning are sensitive to irrelevant prompt changes, even when the problem stays the same. The authors propose SCAPO, an improved training method that assigns credit to individual tokens based on how stable they are under these prompt variations.

trainingreasoningalignment

Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text

Sep 30, 2026

Dulhan Jayalath, Oiwi Parker Jones

Brain-to-text decoders can accidentally learn from word timing patterns instead of brain activity; removing this shortcut with independent window processing makes the task genuinely harder but enables better learning from actual neural signals.

This paper reveals that a major brain-to-text decoding method was exploiting timing shortcuts from word duration patterns rather than learning from actual brain signals. By processing brain windows independently instead of jointly, the authors eliminate this shortcut and achieve better performance (36.6% word error rate) using simpler methods like prediction aggregation and language model priors.

evaluationreasoningapplications

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Sep 30, 2026

Young-Jun Lee, Jinheon Baek, Soyeong Jeong et al.

By treating web search and problem-solving as co-evolving processes rather than separate steps, EvoDuet helps LLMs avoid getting stuck when they need external knowledge, improving discovery performance by 4-21% across scientific optimization tasks.

EvoDuet is a method that improves how AI models search the web while solving scientific problems. It co-evolves search queries and solutions together, letting the model decide when to search for new information versus reusing old documents. The system uses an inner loop to refine searches and an outer loop to generate solutions, achieving significant improvements on optimization tasks.

agentsreasoningapplications

Cogentic: Multi-Agent Orchestration for Automated Proof Discovery

Sep 30, 2026

Yang Cai, Vineet Gupta, Yanchen Jiang et al.

Multi-agent orchestration with verification loops can solve research-level math problems that single-shot generation cannot—by exploring multiple directions, catching errors through adversarial checking, and retaining progress across long reasoning horizons.

Cogentic is a multi-agent system that orchestrates teams of AI provers to tackle open research problems in mathematics and theoretical computer science. Instead of relying on single attempts, it uses an iterative loop where specialized agents explore different proof directions, verify results against each other, and build on confirmed findings stored in a persistent ledger.

agentsreasoningevaluation

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Sep 30, 2026

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang et al.

By separating semantic planning from physical execution and using failure evidence to guide targeted capability improvements, robots can achieve 4x better performance on long-horizon manipulation tasks compared to frozen policies.

DynaHarness is a system that improves robot manipulation by coupling semantic reasoning with physical execution monitoring. It uses a two-level architecture where a 'slow brain' plans high-level actions and a 'fast brain' grounds and monitors execution, refusing unsafe actions and requesting replans when needed. The system learns from failures to improve reusable capabilities.

agentsreasoningtraining

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Sep 30, 2026

Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé

For autonomous ML engineering, a minimal harness giving an LLM direct access to read, write, and bash commands performs as well as complex multi-agent systems—the model itself, not the infrastructure, drives performance.

This paper challenges the complexity of modern ML engineering agents by comparing elaborate multi-agent systems against a simple baseline where an LLM directly accesses code execution tools. The authors find that under equal time budgets, simpler agents perform as well as complex orchestrated systems, suggesting the LLM backbone matters far more than the surrounding machinery.

agentsreasoningarchitecture

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Sep 29, 2026

Jaewoo Jung, Hyeonseo Yu, Honggyu An et al.

Training MLLMs to internally reconstruct 3D scene geometry (even in compact form) improves spatial reasoning and 3D understanding, suggesting that learning to imagine scenes is more effective than explicit geometric supervision alone.

This paper teaches multimodal AI models to reason about 3D scenes by first imagining a compact 3D representation before answering questions. Instead of relying on detailed geometric details, the model learns to assemble a coarse 3D layout from multiple viewpoints—similar to how humans understand 3D space—then uses this mental model to answer spatial reasoning questions more accurately.

multimodalreasoningarchitecture

Breakdown of Local Denoising as Semantic Speciation

Sep 29, 2026

Guangkuo Liu, Mert Okyay, Yifan F. Zhang et al.

Semantic information explains why generative models transition from local to nonlocal generation at nearly the same time—this connection becomes sharper as models scale, suggesting a fundamental phase transition in how semantic structure emerges.

This paper investigates why generative models show two key moments happening nearly simultaneously: when samples commit to a semantic class (speciation) and when local context becomes insufficient for generation (nonlocality). The authors prove these windows are causally linked through semantic information distribution, showing nonlocality must occur within speciation under natural conditions.

reasoningscalingarchitecture

A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization

Sep 29, 2026

Jianru Shen

When evaluating LLM graph reconstruction, a single distance metric masks whether errors come from adding edges, removing edges, or both—the paper provides a mathematical framework to detect mixed editing and reveals different models have fundamentally different failure modes.

This paper analyzes how language models reconstruct graphs, proving mathematical bounds on the Wasserstein distance between original and reconstructed graph spectra. The bounds reveal whether a model only adds edges, only removes them, or does both—information hidden by standard distance metrics.

evaluationreasoning

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Sep 29, 2026

Hui Ren, Lei Fan, Henry Pao et al.

For long-video question answering, visually grounding object identity across time—not just retrieving relevant clips—is critical. GEB shows that organizing observations into entity biographies improves accuracy by 4+ points on day-long and week-long videos.

This paper solves a key challenge in long-video understanding: tracking the same physical object across hours or days of footage. The authors introduce Grounded Entity Biographies (GEB), which groups visual observations of the same object into retrievable "biographies" that preserve context.

multimodalreasoning

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Sep 29, 2026

Dor Tirosh, Ido Amos, Mor Geva

By adding recurrent feedback during pretraining via teacher-supervised state prediction, language models can improve reasoning and task performance without sacrificing training efficiency, suggesting that feed-forward architectures unnecessarily limit information flow.

This paper introduces LIFT, a transformer architecture that enables information to flow backward across layers during language model generation. Instead of the standard feed-forward design, LIFT uses teacher supervision during pretraining to train models to predict both the next token and a dense state representation derived from a teacher model.

architecturetrainingreasoning

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Sep 29, 2026

Paras Dahal, Anton Bakhtin, Taco Cohen et al.

Spending computation on structured decision-making about *how* to solve a problem—not just solving it directly—becomes increasingly valuable as agent tasks scale to longer horizons.

This paper introduces agentic meta-reasoning, a control system that helps AI agents manage long, complex tasks by making explicit decisions about which work to pursue, when to restart, and when to stop. A controller tracks progress compactly and decides next steps, while workers execute the actual task.

agentsreasoning

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Sep 29, 2026

Cheng Qian, Kunlun Zhu, Beibin Li et al.

AI systems can improve other AI systems' performance by learning to build better execution environments—a form of test-time optimization that's reusable across tasks without modifying model weights.

This paper studies how an AI system (Builder) can learn to design better execution environments for another AI system (Target) without changing either model's weights. The Builder learns reusable principles called Meta-Skills from feedback on development tasks, then applies these to construct better environments for new tasks. Results show significant performance improvements across benchmarks.

agentsreasoningtraining

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

Sep 29, 2026

Rishabh Agrawal, Hejie Cui, Shasha Li et al.

Selective learning from feedback—keeping corrections that significantly change model behavior while filtering out those that don't—improves advisor performance and generalization to new tasks and model families.

This paper presents AdviSD, a method for training small AI advisors that guide frozen large language models through natural-language feedback. The key innovation is selectively learning from corrections based on how much they actually change the executor's behavior, avoiding learning from corrections that don't meaningfully affect outcomes.

trainingreasoningefficiency

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Sep 28, 2026

Yijia Fan, Ziqi Huang, Zhongang Cai et al.

Unified models can learn to self-correct their outputs by applying RL to complete reflection loops, where both the reasoning about what's wrong and the actual image fixes improve together without needing external verifiers.

This paper presents UMM-Reflection, a method that teaches unified multimodal models to critique and fix their own image generations through reinforcement learning. Instead of just generating images once, the model can now look at what it created, identify problems, revise the image, and repeat—all within a single model.

trainingreasoningmultimodal

Scaling Long-Form Story Generation via Narrative State Tracking

Sep 28, 2026

Zhennan Wan, Jianfei Chen

Explicitly tracking narrative state (characters, events, plot requirements) as a structured agent task enables LLMs to write coherent long-form stories without special training, scaling from short stories to novel-length works.

This paper introduces NstAgent, a framework that helps large language models write longer stories by tracking narrative elements like characters and plot points. The system maintains consistency across 10K-100K word stories without degrading quality, addressing a key challenge in scaling creative writing to full-length novels.

agentsreasoningapplications

Neural Harmonic Measure Operator

Sep 28, 2026

Jinjin He, Sinan Wang, Yuchen Sun et al.

NHMO enables fast, geometry-aware PDE solving by learning a domain-specific kernel once, then reusing it for any boundary conditions or sources—eliminating expensive retraining for each new problem.

This paper presents Neural Harmonic Measure Operator (NHMO), a neural network approach that solves elliptic PDEs (like Laplace and Poisson equations) on domains with varying shapes.

reasoningefficiencyarchitecture

Improving Test-Time Scaling with Adaptive Looped Transformers

Sep 28, 2026

Yichen You, Tianyu Fu, Aosong Feng et al.

Adaptive token-level iteration selection during inference can significantly improve test-time scaling efficiency—TaH2 achieves 53% better accuracy-per-compute gains than fixed looping by intelligently deciding which tokens deserve extra processing passes.

This paper improves how AI models use extra computation time during inference by introducing TaH2, which selectively applies multiple processing passes to tokens that benefit most from them. Unlike standard looped transformers that process every token multiple times, TaH2 learns which tokens need extra iterations, achieving better accuracy gains per unit of compute on math reasoning tasks.

efficiencyreasoning

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Sep 28, 2026

Jonathan Light, Christopher Zhang Cui, Jeonghye Kim et al.

Agents can learn to act better by learning to explain their actions—training on self-generated retrospections alone improves future performance without RL, suggesting explanation is a useful learning signal for behavior improvement.

This paper shows that language model agents can improve their performance by training on self-generated explanations of their own experiences, without needing reinforcement learning or external rewards. The method, called Retrospection-Only Fine-Tuning (ROFT), has an agent attempt tasks, generate explanations of what happened, and then fine-tune on predicting those explanations.

trainingagentsreasoning

Harness Learning Enables Generalizable Test-Time Adaptation

Sep 28, 2026

Alvin Zhang, Xuecheng Liu, Zixuan Wang et al.

Language model agents can adapt to new tasks by learning to revise their execution harness (program structure) rather than their weights, enabling test-time adaptation that generalizes to unseen tasks.

This paper introduces harness learning, a method where an AI agent learns to improve its own executable program (harness) that controls how a language model makes decisions and uses tools. Instead of changing the model's weights, a separate proposer model learns to revise the harness structure based on task feedback.

agentstrainingreasoning
trainingevaluationreasoning

Trust Guided Decision Transformer

Sep 25, 2026

Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi et al.

A model can detect when its own context has become unreliable by monitoring prediction error—use this self-awareness to filter bad context before making decisions, rather than blindly trusting a critic's action choices.

Decision Transformers struggle on long rollouts because their conditioning context drifts from training data. This paper proposes Trust Guided Decision Transformer (TGDT), which uses the model's own prediction errors to identify when context becomes unreliable, then filters out bad context before selecting actions.

reasoning

Strategically Diverse Sampling for Self-Training

Sep 25, 2026

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

For self-training, sampling diverse problem-solving strategies matters more than correctness or teacher model size—a small model trained on varied approaches beats distillation from a 235B teacher.

This paper shows that self-training works better when you sample diverse problem-solving approaches rather than just correct answers. The authors introduce GROOT (a tree-based sampling method) and Verbalized Sampling to generate strategically different solutions, and find that models trained on diverse but incorrect traces outperform those trained on correct answers from much larger models.

trainingreasoning

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

Sep 25, 2026

Carolina Fortuna, Blaz Bertalanic

Multi-agent LLM scaling isn't automatic—task structure and how you combine outputs matter more than team size. On some tasks, adding agents barely helps because the underlying model's systematic biases affect all team members the same way.

This paper analyzes how multi-agent LLM teams scale based on task structure, using Steiner's taxonomy to distinguish disjunctive tasks (where one correct answer helps) from compensatory tasks (where averaging helps).

agentsscalingreasoning

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

Sep 24, 2026

Jiabin Qiu, Zixuan Chen, Hongye Cao et al.

World models for planning need to preserve action-discriminative information through training, not just minimize prediction error—this simple insight significantly improves both simulation and real-world robotic control performance.

This paper shows that world models trained only to predict what actually happens often fail at model predictive control, which requires comparing different action choices.

reasoningagentsefficiency

Agentic Detection of Online Conspiracies

Sep 24, 2026

Lior Biton, Oren Tsur

Detecting conspiracy theories requires understanding speaker intent through social context, not just analyzing text—and AI agents that adaptively query relevant context perform better than models that process all context at once.

This paper tackles conspiracy detection on social media by recognizing that the same text can express endorsement, criticism, or satire depending on context and speaker intent. The authors propose an agentic framework with tools for querying social context (like user history and network information) to infer whether someone genuinely believes conspiracy theories or is being sarcastic.

agentsreasoningsafety

RAPID: Robot Agentic Programming from Demonstrations

Sep 24, 2026

Yuyao Liu, Jiayuan Mao, David Hsu et al.

By combining visual demonstrations with agentic code refinement and object-centric representations, RAPID enables robots to learn generalizable manipulation skills that transfer across object variations and scene configurations.

RAPID automatically generates reusable robot programs from a single human video demonstration. It uses AI coding agents to create, test, and refine programs that work across different objects and environments by learning the underlying strategy rather than memorizing specific motions.

agentsreasoningapplications

Rolling-WAM: World Action Models with Rolling Imagination

Sep 24, 2026

Yinghua Zhou, Junjie Ye, Yiqi Zhao et al.

Distributing prediction computation across replanning cycles via a rolling noise schedule achieves 4.5x speedup in robot control latency without sacrificing task performance.

Rolling-WAM speeds up robot control by spreading the computation of predicting future actions and images across multiple planning cycles instead of doing it all at once. Instead of fully planning the entire future from scratch each time, it maintains a sliding window of partially-computed predictions at different stages, letting them gradually refine as new camera data arrives.

efficiencyagentsreasoning

Coding Agents for Generalized Task and Motion Planning Problems

Sep 24, 2026

Matteo Merler, Bowen Li, Josh Roy et al.

Coding agents can automatically synthesize generalizable planning programs better than traditional planners or one-shot LLM generation, suggesting LLMs writing code is a practical approach to automating robotics planning.

This paper shows that AI coding agents (like Claude and GPT models) can automatically write programs that solve robot task-and-motion planning problems across different scenarios. Given access to a simulator, agents develop reusable code within a budget, then freeze it for evaluation on new instances.

agentsreasoningapplications

To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

Sep 24, 2026

Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri

Audio language models have a significant blind spot: they verify written facts reliably but fail on identical claims when spoken. Retrieval-augmented approaches help only when combined with explicit reasoning, not retrieval alone.

VeriSpeak is a benchmark for fact-checking spoken claims using audio language models. It contains nearly 4,000 spoken statements about real-world facts and tests whether models can verify claims directly from speech, especially when given retrieved text evidence.

evaluationmultimodalreasoning

PoEM: Predicting RL Outcomes from Existing Policies

Sep 24, 2026

Kimia Hamidieh, Giannis Daras, Antonio Torralba

You can predict RL outcomes for new reward functions by combining existing trained models mathematically, avoiding the computational cost of retraining—useful when experimenting with different objectives or combining multiple goals.

PoEM predicts what a reinforcement learning model will do with a new reward function by combining existing models trained on different rewards, without running expensive RL training. The method works by finding that RL policies live in a low-rank space that can be reconstructed as a linear combination of existing policies.

trainingefficiencyreasoning

Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority

Sep 24, 2026

Mehmet Iscan

Separating candidate generation from verification through an external gate with formal constraints can achieve high safety (zero false releases on benchmark) while maintaining usability, but real-world deployment requires testing with actual users and measuring gate sensitivity.

This paper presents a safety protocol for AI-assisted mechatronic systems that separates candidate generation from release decisions. A frozen 4-billion-parameter language model generates plans, but they're only released when an external verification gate confirms both required facts using a formal grammar.

safetyevaluationreasoning

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Sep 24, 2026

Ming Zhang, Zhenghao Xiang, Peizhong Gao et al.

Current AI systems can learn new rules through exploration, but performance is inconsistent and fragile—gains from exploration can reverse with continued interaction, suggesting exploration capabilities need significant improvement.

ExplorationBench is a benchmark that tests whether AI systems can genuinely explore and discover new rules in unfamiliar environments, rather than just recalling training data. It uses two 'Alien Worlds' with executable rules that differ from real-world knowledge, forcing systems to experiment, form hypotheses, and learn through interaction rather than memorization.

evaluationreasoningagents

Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers

Sep 24, 2026

Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek et al.

Training latent dynamics models for long-horizon stability requires explicitly optimizing for multi-step rollout accuracy, not just reconstruction—this restructures the solution space in ways that conventional metrics don't capture.

This paper shows that neural surrogate models for physics simulations fail during long predictions not because of poor compression, but because they're trained only to reconstruct data. The authors introduce training techniques—including Koopman operator learning and noise injection—that restructure the latent space to support stable long-horizon forecasting.

efficiencytrainingreasoning

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

Sep 24, 2026

Xinyue Zeng, Jiawei Zhang, Yujun Yan et al.

Long-horizon reasoning failures in LLMs stem from structural biases in the reasoning space itself, not just model capacity—and injecting geometric structure into the reasoning process can dramatically improve performance on hard problems.

This paper addresses why large language models struggle with long-horizon reasoning tasks by identifying two key problems: exploration bias (getting stuck in locally plausible but structurally weak paths) and compounding bias (small errors accumulating over many steps).

reasoningarchitectureevaluation

ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints

Sep 24, 2026

Sriram Kannan, Swetha Saseendran, Vishnu Vardhan Reddy Kandi et al.

Structured event graphs from legal documents improve reasoning tasks like classification and QA, but only when you already have the right documents—they don't help with initial retrieval.

ARGUS builds structured Event Knowledge Graphs from employment-discrimination legal complaints using a 5W1H schema and LLMs. The system extracts facts, organizes events with temporal and causal relationships, and merges them into document-level graphs. Testing shows these graphs improve claim classification and help answer legal questions when relevant documents are already retrieved.

dataapplicationsreasoning

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

Sep 24, 2026

Nayoung Choi, Shengjian Chen, Xiaokai Wei et al.

Optimizing search query understanding components individually with search-engine-derived rewards outperforms single end-to-end optimization, showing that understanding how each component affects downstream retrieval matters more than just matching labels.

This paper presents a reinforcement learning framework for query understanding in search systems that optimizes multiple components (like intent classification and query expansion) separately using rewards from live search engine interactions, rather than treating it as a single end-to-end problem.

trainingreasoningapplications

GridSFM: A Foundation Model for Solving AC Optimal Power Flow

Sep 24, 2026

Luke Bhan, Weiwei Yang, Margaret Capetz et al.

Foundation models can solve complex physics optimization problems across different system sizes by combining pretraining on diverse topologies with physics-informed fine-tuning, enabling practical deployment on real power grids without retraining from scratch.

GridSFM is a foundation model that solves AC Optimal Power Flow (a critical power grid optimization problem) by pretraining a 15M-parameter graph neural network across 54 different grid topologies, then fine-tuning it with physics-informed methods.

reasoningapplications

Learning and interpreting policies for simultaneous entanglement requests in quantum networks

Sep 24, 2026

Leon Rode, Sumeet Khatri, Supartha Podder

RL-trained policies can schedule quantum network resources far more efficiently than hand-crafted heuristics, and LLMs can extract interpretable rules from these policies—useful as quantum networks scale beyond what's computationally trainable.

This paper uses reinforcement learning to schedule entanglement resources in quantum networks, enabling multiple quantum tasks to run simultaneously with minimal resources.

agentsreasoning

Does a model's stated reason for rejecting a candidate do any work?

Sep 24, 2026

Archit Rastogi

Language models often cite missing facts when rejecting candidates, but careful testing shows these stated reasons have limited causal influence on their actual choices, raising questions about whether models are genuinely reasoning or post-hoc rationalizing.

This paper tests whether language models' stated reasons for rejecting candidates actually influence their decisions.

evaluationreasoningalignment

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Sep 22, 2026

Zhihao Zhan, Ting Song, Li Dong et al.

Decentralized multi-agent systems can outperform centralized ones by letting workers self-organize through shared workspaces and asynchronous communication, enabling practical scaling to thousands of agents for complex tasks.

Agensh is a multi-agent system that scales to 1,024 agents without needing a central coordinator. Instead of one orchestrator managing all tasks, workers self-organize by sharing a workspace, communicating asynchronously, and claiming their own work.

agentsscalingreasoning

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Sep 22, 2026

Haobo Zheng, Tan Tang, Yan Chen et al.

For multi-party dialogue, tracking speaker identity and relationships separately from content is crucial—this dual-track approach outperforms general-purpose memory systems that try to handle everything at once.

This paper introduces SpeakerMem-R1, a memory system for multi-party conversations that tracks who said what and how people relate to each other. Unlike general LLMs that lose track of speakers and relationships, it uses a dual-track approach: storing exact messages labeled by speaker plus derived relationship states, organized by person and group.

reasoningagentsevaluation

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Sep 22, 2026

Trang Nguyen, Eulrang Cho, Bingqing Chen et al.

By compacting context through selective truncation rather than rephrasing, agents can maintain performance on long-horizon coding tasks while reducing inference costs significantly—making test-time scaling more economical.

CliffCompaction is a technique that compresses long conversation histories for AI coding agents by selectively removing less important content while keeping everything else unchanged. This cuts costs by up to 50% while maintaining performance, enabling agents to work on complex coding problems that span millions of tokens across multiple sessions.

efficiencyagentsreasoning

Diffusion-Induced Spatial Attention Overlapping Community Detection

Sep 22, 2026

Kosti Koistinen, Vesa Kuikka, Joni Herttuainen et al.

By using diffusion-inspired spatial attention instead of local message passing, DISCO better captures long-range dependencies and community boundaries, making it useful for both static community detection and tracking structural anomalies in dynamic networks.

DISCO is a deep learning method for detecting overlapping communities in networks—groups where nodes belong to multiple communities simultaneously. It combines attention mechanisms with diffusion-based structural priors to overcome limitations of standard graph neural networks, and demonstrates practical value in cybersecurity by tracking how network structure changes over time.

architecturereasoning

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Sep 22, 2026

Yuanteng Chen, Zhilei Liu, Peisong Wang et al.

On-policy distillation recovers reasoning capabilities in ultra-low-bit quantized models by training on the model's own generated outputs rather than fixed data, fixing the exposure bias problem that causes long-form reasoning to fail.

This paper tackles a critical problem in quantized language models: when you compress models to very low precision (under 3 bits), they lose the ability to do math and coding tasks because errors compound during long generation.

trainingefficiencyreasoning

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Sep 21, 2026

Yiran Wang, Xingyilang Yin, Junfu Pu et al.

This benchmark reveals that gameplay requires coordinating visual understanding, instruction following, and long-horizon planning—and shows clear performance gaps between model families, providing a standardized way to measure progress on embodied AI tasks.

GameHorizon is a comprehensive dataset and benchmark for evaluating AI models on video game tasks. It includes 5,000 hours of gameplay from 21 AAA games with aligned videos, actions, and instructions at multiple time scales, plus offline and online evaluation tracks to measure how well models understand and execute gameplay across different planning horizons.

evaluationmultimodalreasoning

Critical-State RL: Diagnosing Trainable States for Multi-Turn Tool Use

Sep 21, 2026

Zixiang Chen, Wenting Zhao, Zhepeng Cen et al.

To improve multi-turn tool use, don't train on all failures equally—use diagnostic methods to identify which specific model calls actually control task success, then focus training there.

This paper solves a key problem in multi-turn tool use: identifying which model calls are worth training on. When a task fails, multiple calls could be responsible, but reward signals get muddied by randomness in later steps. The method diagnoses which calls actually control success by separating action-dependent reward changes from downstream noise, then trains only those critical calls.

agentsreasoning

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Sep 21, 2026

Wangbo Yu, Kunhao Liu, Wenbo Hu et al.

By storing observations in a viewpoint-aware implicit 3D memory rather than explicit depth maps, video world models can generate longer, more consistent videos across different camera angles without running out of token budget.

WorldCrafter is a video world model that maintains consistent 3D-aware memory across different camera viewpoints and long time horizons. It uses an implicit memory system that compresses multi-view observations into tokens optimized for the requested viewpoint, enabling realistic minute-long video generation from a single image or text prompt while respecting what was seen before.

architecturemultimodalreasoning

DolphinBench: Mapping the Pareto Frontier of Agent Memory

Sep 21, 2026

Soumil Rathi, Deshraj Yadav, Taranjeet Singh

Most memory benchmarks only measure accuracy, but DolphinBench shows that practical agent memory systems must balance three things: how well they retrieve information, how much they cost, and how fast they respond.

DolphinBench is a benchmark that evaluates how well AI agents use long-term memory to complete real-world tasks, not just answer questions. It includes 500k tokens of realistic user messages per persona and measures accuracy, cost, and latency together—revealing tradeoffs that single-metric benchmarks miss.

evaluationagentsreasoning

Learning Physics from an Imperfect Ancestor

Sep 21, 2026

S. Mohammad Mousavi, Teeratorn Kadeethum, Nikolaos Bouklas et al.

Neural operators don't need to be accurate to be useful—they can guide PINNs away from spurious solutions by providing the right structural prior, enabling reliable PDE solving in regimes where either method alone would fail.

This paper shows how to combine neural operators (fast but inaccurate) with physics-informed neural networks (accurate but optimization-fragile) to solve PDEs reliably. An imperfect neural operator provides a structural hint about which solution the PINN should find, while the PDE residual refines it to high accuracy.

trainingreasoning

Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization

Sep 21, 2026

Filipe Marinho Rocha, Inês Dutra, Vítor Santos Costa et al.

Out-of-distribution generalization requires exact representational equivalence to the generating mechanism, not statistical approximation—a criterion that constrains inference rather than training and explains why neural networks fail on novel entities while logic-based systems succeed.

This paper argues that models generalize beyond their training data only when they compute representations structurally equivalent to the underlying mechanism—not approximations.

reasoningevaluationarchitecture
evaluationreasoning

An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

Sep 18, 2026

Yiming Zhang, Jinghong Zhang, Haoran Zhao et al.

When using RAG with LLMs, blindly trusting all retrieved memories causes hallucinations; a lightweight geometric decision layer can filter unreliable memories without any learned parameters, making RAG safer and more trustworthy.

This paper introduces Memory Decision Layer (MDL), a parameter-free controller that decides whether to trust retrieved memories in RAG systems. It uses three signals—relevance, reliability, and task risk—combined through geometric operations to detect conflicting memories and prevent hallucinations, reducing errors by 56% when memories contradict each other.

reasoningsafetyefficiency

Coding Agents with an Obstacle-Aware Harness for Safe Robot Manipulation

Sep 17, 2026

Bingxin Xu, Yuzhang Shang, Zhen Dong et al.

Language models can understand safety instructions but don't prioritize them during execution—adding explicit obstacle-aware planning and verification mechanisms dramatically improves both task success and collision avoidance in robot control.

This paper addresses safety in coding agents for robot manipulation by introducing SafeHarness, a system that prevents collisions with obstacles during task execution. The key insight is that language models can reason about obstacles but fail to prioritize safety constraints during planning and contact execution.

agentssafetyreasoning

JEPA-Anything: Learning Predictive Models across Different Worlds

Sep 17, 2026

Taoyong Cui, Zhongyao Wang, Xinyue Xu et al.

A single learning principle based on factorized predictions can work across radically different domains, suggesting world modeling doesn't need domain-specific architectures—and can even guide real scientific discovery.

JEPA-Anything is a unified framework for building predictive models across completely different domains—from videos to molecules to weather—using a technique called orthogonal predictive factorization. Instead of training separate models for each domain, it learns to decompose predictions into independent factors that work across vision, biology, physics, and clinical data.

architecturereasoningmultimodal

PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers

Sep 17, 2026

Jiachen Yao, Zi-Siang Hsu, Xi Deng et al.

Existing evaluations of generative inverse solvers miss critical failures like mode collapse and overconfident uncertainty—PosteriorBench reveals these gaps by directly comparing predicted solution distributions to ground-truth posteriors.

PosteriorBench is a benchmark for evaluating how well generative models solve inverse problems by checking if they capture the full range of possible solutions, not just single best guesses. It tests four physics problems using reference posteriors from MCMC and rejection sampling, with metrics measuring accuracy, uncertainty, and distributional fit.

evaluationreasoning