ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2567 papers38 this month12 topics
AllTraining 46Reasoning 42Efficiency 33Evaluation 32Agents 27Architecture 18Multimodal 17Applications 16Data 9Safety 9scaling 5Alignment 2

Oct 5 – Oct 11(8)

Base Models Can Reason By Taking a Cue From Training Data

Oct 5, 2026

Sophie L. Wang, Amil Dravid, Rulin Shao et al.

Base models already contain reasoning capabilities encoded in their training data—you can unlock them by conditioning on the right token cues, without needing expensive RL fine-tuning.

This paper shows that base language models can achieve reasoning performance comparable to RL-trained models by using specific starting tokens (like "Okay" or "Alright") that trigger learned associations from training data. The authors demonstrate they can create new reasoning cues through data interventions and trace these effects back to specific document types in the training set.

trainingreasoningdata

BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance

Oct 5, 2026

Haojin Deng, Zhiping Lin, Yimin Yang

Monitoring centroid geometry during training can help detect and reduce spurious feature reliance, but attribute information remains partially recoverable—suggesting regularization alone isn't sufficient for complete bias removal.

BiasFlow is a monitoring toolkit that tracks how neural network backbones rely on spurious features (like gender in face recognition) through geometric analysis of feature centroids.

Sep 28 – Oct 4(48)

What Should World Models Forget? Stratified Retention for Continual Adaptation

Oct 2, 2026

Nishit Anand, Ramani Duraiswami, Dinesh Manocha

World models need stratified forgetting strategies that preserve physical invariants while quickly adapting to environmental changes—standard continual learning metrics fail to capture this distinction and incorrectly reward frozen models.

This paper addresses a fundamental problem in continual learning for world models: knowing what to forget. Unlike traditional learning where correct labels stay correct, world models operate in changing environments where outdated knowledge must be discarded.

trainingevaluationreasoning

From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing

Oct 2, 2026

Kuangyu Ding, Gesualdo Scutari

Graph decomposition into tree blocks enables more efficient decentralized optimization by jointly designing subproblems and communication patterns, with convergence rates that explicitly depend on network topology and function properties.

This paper develops a new framework for distributed optimization over networks where agents minimize functions while only communicating with neighbors. Instead of traditional mixing-based approaches, the method decomposes the network graph into tree-structured blocks, with agents cooperatively solving subproblems via message passing.

Sep 21 – Sep 27(19)

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

Sep 25, 2026

Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe et al.

Training language models to predict their confidence in intermediate reasoning steps—using only self-supervised learning—makes them generate shorter reasoning traces at inference time without any explicit length penalties or early-stopping mechanisms.

This paper shows that reasoning models can generate shorter, more efficient reasoning traces by learning to predict their own confidence in answers—without explicitly optimizing for length.

trainingefficiencyreasoning

First-Order Stationarity of Reverse Diffusions

Sep 25, 2026

Zhifeng Chen, Chenyang Jiang, Yazhen Wang

Diffusion models have provable convergence guarantees similar to optimization algorithms—reverse diffusions contract divergence exponentially fast, and discrete samplers achieve measurable stationarity bounds that don't depend on data convexity.

This paper connects optimization theory to diffusion models by proving that reverse-time diffusion processes contract Fisher divergence at exponential rates under strong convexity conditions. The authors also establish first-order stationarity bounds for practical discrete samplers, showing how optimization guarantees translate to sampling quality without requiring global convexity.

Sep 14 – Sep 20(25)

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Sep 18, 2026

Hongyang Du, Lan Yan, Christian Flores et al.

Procedural memory—a continuously updated library of natural-language design skills—enables frozen frontier models to improve at complex agentic tasks by learning from execution failures without model retraining or human annotation.

This paper shows how a frozen AI model can continuously improve at graphic design by building and refining a library of reusable design procedures from real user projects. Without updating the model's weights or using human labels, the system learns 139 design skills from 1,406 real briefs, improving success rates from 73% to 99% by accumulating new procedures and fixing failed ones.

agentstrainingapplications

Cross-sector generalization of accident-process role classification in occupational accident narratives

Sep 18, 2026

Aho Yapi, Pierre Latouche, Arnaud Guillin et al.

Fine-tuned language models can generalize accident report classification across different industries without retraining, enabling scalable occupational safety analysis across sectors.

This paper develops an automated system to classify key information in occupational accident reports (work situations, unsafe conditions, events, consequences) and tests whether models trained on construction-sector narratives can work across different industries like metallurgy and chemistry.

safetyevaluationtraining

Learning to Read the Contextual Tokens in Diffusion Transformers

Oct 5, 2026

Omer Dahary, Etai Sella, Hadar Averbuch-Elor et al.

Text tokens in image-generating transformers develop interpretable semantic representations of the emerging image that can be read with an LLM probe—and explicitly training to strengthen these representations improves generation quality.

This paper reveals what text tokens learn during image generation in multimodal diffusion transformers. Researchers built a tool to 'read' these hidden representations by connecting them to a language model, discovering they encode rich scene information early in generation. They then used these insights to improve image quality through a new training technique.

multimodaltraining

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Oct 5, 2026

Benhao Huang, Chufan Shi, Junlin Chen et al.

Looped models can be dramatically more efficient by recognizing that recurrent states converge to fixed points, enabling truncated backpropagation, KV cache sharing, and faster RL—with learned depth priors and orthogonal injection providing better supervision than existing approaches.

This paper optimizes looped language models—models that process information through multiple recurrent passes—by leveraging fixed points in recurrent states.

trainingefficiencyarchitecture

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Oct 5, 2026

Haozhen Zhang, Haodong Yue, Quanyu Long et al.

Instead of pre-processing all memory upfront, MemPilot learns to make runtime decisions about memory curation, letting developers trade off accuracy against computational cost and speed based on their needs.

MemPilot is a framework that helps LLM agents manage memory more efficiently by deciding when to retrieve pre-stored information versus when to process raw conversation history on-demand.

agentsefficiencytraining

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Oct 5, 2026

Yifan Zhang, Yutong Dai, Viraj Prabhu et al.

Self-verification through conformal methods lets web agents learn from their own reasoning about task progress, eliminating the need for expensive judge calls at deployment while improving training efficiency.

CLIFT trains web agents to complete browser tasks by having them verify their own actions through natural-language questions, creating a reusable signal that works both during training (with sparse rewards) and at test time (without expensive external judges). The method achieves state-of-the-art results on multiple web agent benchmarks and transfers across different models.

trainingagentsreasoning

Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors

Oct 5, 2026

Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi et al.

You can estimate sleep heart rate from accelerometer motion signals using deep learning, trading off some accuracy for broader coverage—useful for extracting cardiac insights from existing wearable data without optical sensors.

This paper presents SeqSmoother, a transformer-based model that estimates heart rate during sleep using only wrist accelerometer data, without requiring optical sensors.

trainingevaluationapplications

Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Oct 5, 2026

Sahil Mahendrakar

Knowledge distillation can compress speech synthesis models by 10x with minimal quality loss by separating the text-to-features and features-to-audio tasks and training them independently against a frozen teacher.

Paradee is a tiny text-to-speech model created by distilling a larger 82M-parameter teacher into just 8M parameters while keeping the same voice quality. The authors use a two-stage training approach: first synthesizing training data with the teacher model, then training separate text and audio components before combining them.

efficiencytrainingarchitecture
trainingefficiency

LESSER: Post-Training Data Selection with Output-Layer Gradients

Oct 2, 2026

Lyuxin David Zhang, Eric Wong, Surbhi Goel et al.

You can select effective training data for LLMs using only output-layer gradients from forward passes, cutting computation costs by ~10× compared to full-gradient methods without sacrificing downstream performance.

This paper proposes LESSER, a method for selecting high-quality training data for large language models by using only output-layer gradients instead of full-parameter gradients. By leveraging cheaper forward passes rather than expensive backward passes, LESSER reduces computational cost by 9.7× for supervised fine-tuning while maintaining performance comparable to full-gradient selection methods.

trainingefficiencydata

Language Models that Play Chess and Explain Their Moves

Oct 2, 2026

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Language models can match expert-level performance in specialized domains by distilling knowledge from silent expert systems through iterative refinement, opening a path to explainable AI in games, robotics, and other domains with strong baseline models.

This paper presents Queen, a 4-billion-parameter chess model that combines a silent chess engine with a language model to play at Grandmaster level while explaining its moves.

reasoningtrainingapplications

Planning to Learn

Oct 2, 2026

Ian Osband

Cross-entropy beats policy gradients in classification because it's 'patient'—it optimizes for total error reduction across all future steps, not just immediate accuracy. A simple horizon-aware loss can capture this benefit while staying closer to principled gradient methods.

This paper reveals why cross-entropy outperforms exact policy gradients in classification despite having access to the true label. The key insight is that cross-entropy implicitly accounts for future learning steps, while exact policy gradients are myopic.

trainingreasoning

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Oct 2, 2026

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi et al.

By identifying and selectively training on high-impact token decisions rather than full sequences, you can make diffusion language models learn more efficiently with less data.

This paper introduces Pivot-SD, a training method for masked diffusion language models that focuses on the most impactful decisions during text generation. Instead of training on entire sequences, it identifies 'pivot' tokens—commitments that significantly reduce uncertainty about remaining words—and trains only on those, using success/failure signals to guide learning.

trainingefficiencyreasoning

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Oct 2, 2026

Vladislav Gromadskii, David Li, Samson Gourevitch et al.

You can fine-tune fast diffusion generators for reward optimization without expensive reference rollouts by replacing intractable KL penalties with inverse-distillation regularization that provably bounds divergence.

IDRF is a method for fine-tuning masked discrete diffusion models (which generate sequences iteratively by predicting multiple tokens at once) to maximize rewards while staying close to a reference model. Instead of computing intractable likelihood penalties, it uses a clever regularization trick called inverse-distillation that upper-bounds the KL divergence.

trainingefficiencyreasoning

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Oct 1, 2026

Yen-Jen Wang, Haozhe Jiang, Shuying Deng et al.

Robots can improve their own performance through autonomous practice and skill refinement in simulation without updating model weights, then transfer successfully to real hardware—a practical path to reliable robot systems.

RPG is a framework that improves robot performance without retraining models by identifying skills from offline data, practicing in simulation with failure diagnosis, and refining symbolic skills and system prompts.

agentsreasoningtraining

Embedding Prediction Helps Image Generation

Oct 1, 2026

Sihan Xu, Ji Xie, Zilin Wang et al.

Dynamically predicting and updating conditioning embeddings at each generation step improves diffusion model efficiency and quality—you don't need to reuse the same embedding throughout the entire denoising process.

This paper proposes using predicted embeddings as dynamic conditioning signals in diffusion transformers instead of static embeddings. A separate transformer (NEPA) predicts image embeddings at each denoising step, allowing the conditioning to adapt to the current noise level. The approach achieves competitive image generation quality on ImageNet with significantly less training compute.

architectureefficiencytraining

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Oct 1, 2026

Jichao Jiang, Cristian McGee, El Houcine Bergou et al.

TACO cuts optimizer memory from 27.7 GB to 0.16 GB on 13B models by using a sparse, low-precision update strategy based on column-wise signs—making full-parameter fine-tuning practical on consumer GPUs without sacrificing model quality.

TACO is a new optimizer for fine-tuning large language models that dramatically reduces memory usage by storing only tiny gradient components per column instead of full optimizer state. It achieves 174× memory reduction compared to AdamW while maintaining accuracy, enabling fine-tuning of 30-32B models on a single GPU.

efficiencytraining

FERPO: Forward Entropy-Regularized Policy Optimization

Oct 1, 2026

Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

Using forward-KL instead of reverse-KL for policy fitting encourages broader exploration of high-value actions and avoids the computational cost of differentiating critics, leading to faster and more sample-efficient learning.

FERPO is a reinforcement learning algorithm that improves policies by using critic values directly rather than differentiating through them. It derives optimal target actions using entropy regularization and fits the actor to these targets using forward-KL divergence, which encourages exploring multiple high-value action modes while keeping importance weights stable.

trainingefficiencyreasoning

Hierarchical Continuous Diffusion Language Models

Oct 1, 2026

Hui Ren, Zihan Li, Chang Liu et al.

HC-DLM bridges discrete and continuous diffusion by coupling token generation with a shared latent trajectory, enabling better reasoning and constraint satisfaction than purely discrete or continuous approaches.

This paper proposes Hierarchical Continuous Diffusion Language Models (HC-DLM), which combines discrete token generation with continuous latent states in a single denoising process. Unlike existing approaches that treat these separately, HC-DLM uses the continuous latent as the only persistent state, reading tokens from it at each step.

architecturereasoningtraining

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Oct 1, 2026

Shuo Xing, Zilin Dai, Chengyuan Qian et al.

Mathematical reasoning in LLMs isn't a single skill but four distinct capabilities; focusing training on the 'Discovery' bottleneck (finding the right solution strategy) is more effective than generic math training.

This paper diagnoses why LLMs struggle with math by breaking down mathematical reasoning into four components (Discovery, Generation, Digestion, Execution) and shows that Discovery—finding the right approach—is the main bottleneck. The authors then propose a training method that uses these insights to improve math performance across different model sizes.

reasoningtrainingevaluation

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Oct 1, 2026

Cristian McGee, El Houcine Bergou, Aritra Dutta

ZFO decouples direction selection from step-size selection in LLM fine-tuning, using gradient information plus two function evaluations to adaptively choose step sizes that often outperform fixed-step methods without the cost of full line searches.

This paper proposes ZFO, a lightweight optimization framework for fine-tuning large language models that intelligently selects step sizes by combining first-order gradient information with minimal zeroth-order function evaluations.

trainingefficiency

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Oct 1, 2026

Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.

Sparse autoencoder features combined with reinforcement learning enable interpretable, composable control over protein sequence generation—activating 3.75x more targeted features than previous steering approaches and improving predicted biological function.

IDiom is a specialized protein language model trained on 54 million intrinsically disordered protein regions (IDRs) that can generate functional sequences with precise control over biological features.

trainingapplicationsreasoning

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Oct 1, 2026

Zhengming Yu, Junkun Yuan, Haotian Yang et al.

By framing distribution matching as a classification problem with discriminators, DMAD eliminates the memory overhead of auxiliary models while maintaining or improving generation quality—enabling practical few-step visual generation.

DMAD improves fast image and video generation by training lightweight student models to match teacher distributions without needing an auxiliary model. It uses two discriminator heads to learn density ratios directly, making the process more efficient while achieving state-of-the-art quality in one-to-four-step generation across images and videos.

efficiencytrainingarchitecture

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Oct 1, 2026

Joohwan Ko, Tetiana Parshakova, Diana Cai et al.

Quasi-Newton methods—traditionally limited to convex optimization—can now scale to deep learning by using variational objectives to ensure positive curvature and GPU-friendly matrix operations, outperforming Adam on ill-conditioned problems.

SoftServe is a new quasi-Newton optimization method for training deep neural networks that handles the challenges of non-convex optimization and massive parameter counts.

trainingefficiency

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Oct 1, 2026

Siqi Zhu, Suozhi Huang, Kaixuan Zhang et al.

When combining multiple RL-trained teachers into one student, the averaging method and optimizer choice matter more than raw gradient differences—response length weighting and precision loss can swing task performance by 2-5 percentage points.

This paper investigates how multiple teacher models transfer knowledge to a student model during on-policy distillation. The researchers found that loss averaging implicitly weights responses, Adam's optimizer smooths gradient differences, and low-precision arithmetic (BF16) masks small weight updates—with these factors significantly affecting which tasks the student learns best.

trainingefficiency

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Oct 1, 2026

Xuan Zhang, Longtao Zheng, Cunxiao Du et al.

Teaching agents to manage their own context through learned compaction decisions—rather than just handling overflow—improves performance on long-horizon coding tasks by 5-9% across different context window sizes.

AutoCompact trains coding agents to automatically decide when and how to compress their working context during long software engineering tasks. By learning when to discard stale exploration and what state to preserve, the agent improves its ability to solve repository-level coding problems while staying within context limits.

agentsreasoningtraining

When Do Intrinsic Rewards Lead to Exploration?

Oct 1, 2026

Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett

Intrinsic rewards designed to encourage exploration can fail to find the most informative experiences—you need to explicitly measure whether an agent's history can substitute for real experience under different policies.

This paper examines when intrinsic reward signals (like prediction error or curiosity) actually lead to good exploration in reinforcement learning. The authors show that maximizing these rewards doesn't always produce the most informative experiences, propose a formal criterion based on counterfactual information, and demonstrate failures of existing methods with concrete examples.

trainingreasoning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Oct 1, 2026

Lucheng Fu, Kejing Xia, Yiyang Wang et al.

LLM agents can significantly improve performance on knowledge-intensive tasks by learning persistent, source-specific models that evolve through repeated interaction—achieving up to 22.6 point gains over standard retrieval methods.

This paper introduces SourceLearn, a method for LLM agents to develop persistent, reusable understanding of external knowledge sources through repeated interaction.

trainingagentsreasoning

Faynt: Scaling and Optimizing Policies for Competitive Melee

Oct 1, 2026

Ali Janati, Nikita Kuzmin, Rohit Swamy et al.

Scaling, architecture choices, and post-training curricula matter more than model size alone—a smaller, optimized policy outperforms a larger pretrained one, suggesting careful design beats raw parameter count for complex game-playing tasks.

Faynt introduces transformer-based AI policies for Super Smash Bros. Melee that control all 26 characters with a single model. A 10M-parameter version wins 98.4% of same-character matches against existing AI opponents and beats a zero-delay competitor, using techniques like supervised pretraining on 840k human replays, curriculum learning, and distillation from a larger 75M model.

trainingreasoningagents

Finetuning with Sampling: SFT Learns Better Than You Think

Oct 1, 2026

Aayush Karan, Sitan Chen, Yilun Du

SFT isn't inherently worse than RL for posttraining—the gap comes from data distribution mismatch. By reshaping data to be more on-policy before training, SFT can generalize better and forget less than strong RL baselines.

This paper shows that supervised finetuning (SFT) can match or beat reinforcement learning for posttraining if you transform the training data first. The authors use an MCMC sampling algorithm to gradually shift off-policy expert demonstrations toward on-policy trajectories that a reference model can actually learn from.

trainingdata

Sample complexity bounds for categorical Markov random fields via Discrete Diffusions

Oct 1, 2026

Shivam Kumar, Nabarun Deb

Discrete diffusions can sample from structured categorical data with provable sample complexity guarantees, and weight-sharing neural networks learn scores more efficiently than fully-connected ones for this task.

This paper develops theoretical guarantees for sampling from high-dimensional categorical distributions using discrete diffusion models.

trainingreasoningscaling

Local Support Learning

Oct 1, 2026

Assaf Ben-Kish, Akarsh Kumar, James Glass et al.

You can prevent large models from forgetting old skills during new training by using a learnable gate that only activates weight updates when the input matches the current training distribution—no need to store old data.

This paper addresses catastrophic forgetting in large language models by treating it as a geometric problem in weight space.

trainingefficiencyalignment

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Oct 1, 2026

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc et al.

Spatially grounded self-distillation with synthetic data can teach multimodal models better visual reasoning that transfers to real-world tasks, without requiring human annotations or external teachers.

This paper improves multimodal AI models by having them learn from a smarter version of themselves that receives spatial hints about where to look in images. Using synthetic scenes with automatic object labels, the approach trains models to understand spatial relationships without human annotation, and surprisingly, these improvements transfer to real-world vision tasks.

trainingmultimodaldata

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

Oct 1, 2026

Zilin Du, Bowen Yang, Boyang Albert Li

When scaling data selection with neural networks, the standard loss function causes poor generalization—a new loss function (PVM) that matches predicted values pointwise solves this and transfers better across datasets and model scales.

This paper addresses data selection for training large language models by proposing TESS, a framework that uses a neural network to score and select training examples. Unlike existing meta-learning approaches that assign per-sample weights, TESS uses a novel loss function (Pointwise Value Matching) that avoids optimization instability and improves generalization to new datasets and model sizes.

trainingdataefficiency

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

Oct 1, 2026

Yinheng Li, Justin Wagle

Modern LLMs are naturally good at making structured decisions from predefined options without fine-tuning, but targeted fine-tuning helps weaker models and specific tasks like routing—without degrading their conversational abilities.

This paper shows that large language models can already make categorical decisions (choosing from predefined options) without generating text, using their built-in token probabilities. The authors present LLM2Jev, a method to extract these decisions directly and optionally fine-tune models to improve decision-making on specific tasks, while keeping the model's text generation abilities intact.

trainingefficiencyapplications

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

Oct 1, 2026

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

HAO enables stable reinforcement learning under FHE encryption by preventing polynomial approximation errors from accumulating—achieving 0% boundary violations versus 83.8% for unprotected baselines, making privacy-preserving RL practically viable.

This paper solves a critical problem in privacy-preserving reinforcement learning: when you encrypt data with Fully Homomorphic Encryption (FHE) for cloud computation, you must replace nonlinear operations with polynomial approximations, which causes training to diverge.

safetyefficiencytraining

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

Oct 1, 2026

Arman Behnam, Binghui Wang

Current memory systems can't measure the value of memories that are never retrieved.

Memory-augmented language models struggle to identify which memories are actually useful because some memories are never retrieved, making their value impossible to measure.

reasoningevaluationtraining

Semifactual Credit-Augmented Policy Optimization

Sep 30, 2026

Junshu Pan, Zhizhang Fu, Shulin Huang et al.

Token-level stability under prompt variations is a useful training signal for improving reasoning in LLMs—you can boost performance by penalizing tokens that change meaning when irrelevant prompt details change.

This paper identifies that language models trained with reinforcement learning are sensitive to irrelevant prompt changes, even when the problem stays the same. The authors propose SCAPO, an improved training method that assigns credit to individual tokens based on how stable they are under these prompt variations.

trainingreasoningalignment

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Sep 30, 2026

Razan El Mais, Ali Chehab, Ibrahim Issa et al.

When fine-tuning LLMs with differential privacy, untying input/output embeddings outperforms the standard weight-tied design and enables 60% memory savings—suggesting privacy-preserving training requires rethinking standard model architectures.

This paper investigates weight tying (sharing parameters between input and output embeddings) in large language models trained with differential privacy. The authors find that untying embeddings actually improves performance under DP-SGD, achieving up to 4.74% accuracy gains, while also enabling more memory-efficient privacy techniques.

safetyefficiencytraining

Turbo Harness: Instance-Adaptive Harness Optimization

Sep 30, 2026

Tunyu Zhang, Hao Wang, Kai Xu et al.

Adapting execution harnesses to individual task instances—rather than using a single global harness—consistently improves agent performance, and this adaptation can be automated by learning from previous optimization runs.

This paper introduces Turbo Harness, a system that automatically customizes AI agent execution frameworks (harnesses) for individual tasks by learning from past optimization runs. Instead of using one fixed harness for all tasks, it generates task-specific modifications that improve agent performance across diverse domains like interactive tasks, coding, and long-horizon planning.

agentstrainingefficiency

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Sep 30, 2026

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang et al.

By separating semantic planning from physical execution and using failure evidence to guide targeted capability improvements, robots can achieve 4x better performance on long-horizon manipulation tasks compared to frozen policies.

DynaHarness is a system that improves robot manipulation by coupling semantic reasoning with physical execution monitoring. It uses a two-level architecture where a 'slow brain' plans high-level actions and a 'fast brain' grounds and monitors execution, refusing unsafe actions and requesting replans when needed. The system learns from failures to improve reusable capabilities.

agentsreasoningtraining

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Sep 29, 2026

Yi Pan, Haocheng Xi, Kan Zhu et al.

You can compress linear attention's recurrent state to 8-bit without significant quality loss by quantizing only at window boundaries and preserving outliers as special tokens—enabling faster inference on long sequences.

LeapQuant reduces the inference cost of linear attention models by quantizing their recurrent state to 8-bit precision while maintaining accuracy. It uses per-window quantization to limit error buildup and compensator tokens to handle outliers, achieving 2-3.7x speedups on real hardware.

efficiencytraining

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Sep 29, 2026

Kuan-Po Huang, Haohe Liu, Puyuan Peng et al.

Emotion vectors in TTS models can be decomposed into a neutral-shift component and an emotion-specific component—controlling them separately via steering achieves much better emotion control than treating them as a single direction.

This paper improves emotional speech generation by decomposing emotion vectors into shared and residual components, then controlling them separately without retraining the model. The method, EmoRES, significantly outperforms prior vector steering approaches on multiple emotion metrics and human evaluation.

efficiencymultimodaltraining

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Sep 29, 2026

Dor Tirosh, Ido Amos, Mor Geva

By adding recurrent feedback during pretraining via teacher-supervised state prediction, language models can improve reasoning and task performance without sacrificing training efficiency, suggesting that feed-forward architectures unnecessarily limit information flow.

This paper introduces LIFT, a transformer architecture that enables information to flow backward across layers during language model generation. Instead of the standard feed-forward design, LIFT uses teacher supervision during pretraining to train models to predict both the next token and a dense state representation derived from a teacher model.

architecturetrainingreasoning

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Sep 29, 2026

Cheng Qian, Kunlun Zhu, Beibin Li et al.

AI systems can improve other AI systems' performance by learning to build better execution environments—a form of test-time optimization that's reusable across tasks without modifying model weights.

This paper studies how an AI system (Builder) can learn to design better execution environments for another AI system (Target) without changing either model's weights. The Builder learns reusable principles called Meta-Skills from feedback on development tasks, then applies these to construct better environments for new tasks. Results show significant performance improvements across benchmarks.

agentsreasoningtraining

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

Sep 29, 2026

Rishabh Agrawal, Hejie Cui, Shasha Li et al.

Selective learning from feedback—keeping corrections that significantly change model behavior while filtering out those that don't—improves advisor performance and generalization to new tasks and model families.

This paper presents AdviSD, a method for training small AI advisors that guide frozen large language models through natural-language feedback. The key innovation is selectively learning from corrections based on how much they actually change the executor's behavior, avoiding learning from corrections that don't meaningfully affect outcomes.

trainingreasoningefficiency

Telescopic Language Models

Sep 28, 2026

Zhilin Guo, Boqiao Zhang, Hakan Aktas et al.

You can train one model that works well at any depth by supervising random layer prefixes during training—no architectural tricks needed, just a smarter training objective that makes the model elastic across compute budgets.

This paper introduces Telescopic Language Models (TLMs), which train a single model that works effectively at every layer depth rather than requiring separate models for different compute budgets.

trainingefficiencyarchitecture

PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

Sep 28, 2026

Zimo Wang, Junkun Yuan, Angtian Wang et al.

Critic error accumulation is a fundamental bottleneck in distilling video diffusion models; filtering it via projection dramatically improves sample quality without architectural changes or extra computation.

This paper improves video diffusion model distillation by fixing a key problem: critic errors that accumulate during training and degrade sample quality. PDMD uses a simple mathematical projection to filter out these errors while preserving useful learning signals, achieving better video quality with fewer computational steps—all with just a one-line code change.

efficiencytrainingevaluation

Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning

Sep 28, 2026

Yijia Fan, Ziqi Huang, Zhongang Cai et al.

Unified models can learn to self-correct their outputs by applying RL to complete reflection loops, where both the reasoning about what's wrong and the actual image fixes improve together without needing external verifiers.

This paper presents UMM-Reflection, a method that teaches unified multimodal models to critique and fix their own image generations through reinforcement learning. Instead of just generating images once, the model can now look at what it created, identify problems, revise the image, and repeat—all within a single model.

trainingreasoningmultimodal

How to Loop MoE: Flatten the Experts, Untie the Attention

Sep 28, 2026

Shouren Wang, Chuang Ma, Mohsen Hariri et al.

To build better looped MoE models, flatten the expert hierarchy (more experts per layer, more passes) and give each pass independent attention—this lets tokens access more experts while maintaining computational efficiency.

This paper improves looped mixture-of-experts (MoE) models by flattening the expert structure and untying attention parameters. The key insight is that by doubling experts per layer and doubling passes through the network while keeping compute fixed, models can route tokens through more diverse experts, improving performance.

architectureefficiencytraining

KV-streams for Efficient Compaction in Agentic Reinforcement Learning

Sep 28, 2026

Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda et al.

KV-streams enables efficient scaling of agentic LLMs to longer horizons by streaming cached computations rather than recomputing them, making long-context RL training practical without sacrificing performance.

This paper introduces KV-streams, a technique that speeds up training of long-horizon agentic language models by streaming the key-value cache forward during context compaction instead of repeatedly refilling it. The method achieves 2.6-5x training speedup while maintaining performance, and shows that the streamed cache can retain information beyond the visible context window.

efficiencytrainingagents

Towards Communication-Efficient Social Intelligence in Language Agents

Sep 28, 2026

Linxiao Gong, Yijie Xu, Tianfu Wang et al.

TACT enables language agents to achieve better social outcomes while using fewer tokens and messages by having specialized teachers refine communication strategy and expression, then distilling improvements into the student agent.

This paper introduces Teacher-Assisted Communication Training (TACT), a method that helps language agents communicate more efficiently during social interactions. TACT improves how agents negotiate and coordinate by having a teacher refine both what agents say (expression) and how they say it (strategy), then distills these improvements back into the student agent.

trainingagentsefficiency

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Sep 28, 2026

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan et al.

When adapting pre-trained vision models to use linear attention, directly copy MLP weights but distill attention behavior—this simple strategy closes the performance gap between efficient and standard transformers.

This paper shows how to initialize linear Vision Transformers (efficient attention models) using weights from standard Softmax ViTs. The key insight: copy the MLP layers directly since they learn general representations, but use distillation to transfer the attention mechanism since it's operator-specific.

efficiencyarchitecturetraining

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

Sep 28, 2026

Jonathan Light, Christopher Zhang Cui, Jeonghye Kim et al.

Agents can learn to act better by learning to explain their actions—training on self-generated retrospections alone improves future performance without RL, suggesting explanation is a useful learning signal for behavior improvement.

This paper shows that language model agents can improve their performance by training on self-generated explanations of their own experiences, without needing reinforcement learning or external rewards. The method, called Retrospection-Only Fine-Tuning (ROFT), has an agent attempt tasks, generate explanations of what happened, and then fine-tune on predicting those explanations.

trainingagentsreasoning

Harness Learning Enables Generalizable Test-Time Adaptation

Sep 28, 2026

Alvin Zhang, Xuecheng Liu, Zixuan Wang et al.

Language model agents can adapt to new tasks by learning to revise their execution harness (program structure) rather than their weights, enabling test-time adaptation that generalizes to unseen tasks.

This paper introduces harness learning, a method where an AI agent learns to improve its own executable program (harness) that controls how a language model makes decisions and uses tools. Instead of changing the model's weights, a separate proposer model learns to revise the harness structure based on task feedback.

agentstrainingreasoning
trainingevaluationreasoning

New LoRA Skills Should Read but Never Write

Sep 25, 2026

Zeyan Li, Panqi Yang, Qirong Guo et al.

When combining multiple LoRA adapters, the internal representation and directional coupling between them matters more than the adapter weights themselves—fixing these choices lets you add skills sequentially without degrading previous ones.

This paper solves the problem of combining multiple fine-tuned LoRA adapters into a single model without interference. The key insight is that LoRA updates have multiple equivalent forms, and the choice matters when combining adapters.

trainingefficiency

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Sep 25, 2026

Varun Reddy, Bernardo L. Sabatini, Houman Safaai

Common-mode error in direct feedback alignment causes training stalls by saturating hidden units; this can be prevented by centering batch errors or calibrating the readout baseline, enabling faster learning without changing the core algorithm.

Direct feedback alignment trains neural networks using fixed random error projections, but gets stuck learning near a baseline predictor. The paper identifies that a shared error component across inputs drives hidden units to saturation, slowing learning. Simple fixes like centering errors or adjusting the baseline readout can prevent this collapse and speed up training.

trainingefficiency

Strategically Diverse Sampling for Self-Training

Sep 25, 2026

Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

For self-training, sampling diverse problem-solving strategies matters more than correctness or teacher model size—a small model trained on varied approaches beats distillation from a 235B teacher.

This paper shows that self-training works better when you sample diverse problem-solving approaches rather than just correct answers. The authors introduce GROOT (a tree-based sampling method) and Verbalized Sampling to generate strategically different solutions, and find that models trained on diverse but incorrect traces outperform those trained on correct answers from much larger models.

trainingreasoning

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

Sep 24, 2026

Wenhao Li, Zhibin Wu, Chong Xiao et al.

Using LLM-generated semantics as a shared anchor point for aligning incomplete multimodal data is more robust than trying to reconstruct missing modalities or design complex fusion mechanisms.

SemMSA tackles multimodal sentiment analysis when some data is missing by using large language models to create rich semantic representations that ground all modalities together. Instead of reconstructing missing features, it aligns visual, acoustic, and text representations through spectral methods, achieving better results on standard benchmarks.

multimodalalignmenttraining

PoEM: Predicting RL Outcomes from Existing Policies

Sep 24, 2026

Kimia Hamidieh, Giannis Daras, Antonio Torralba

You can predict RL outcomes for new reward functions by combining existing trained models mathematically, avoiding the computational cost of retraining—useful when experimenting with different objectives or combining multiple goals.

PoEM predicts what a reinforcement learning model will do with a new reward function by combining existing models trained on different rewards, without running expensive RL training. The method works by finding that RL policies live in a low-rank space that can be reconstructed as a linear combination of existing policies.

trainingefficiencyreasoning

Anchored Extra-Proximal Methods: Optimal Higher-Order Methods for Monotone Inclusion Problems

Sep 24, 2026

Ruichen Jiang, TaeHo Yoon

Higher-order optimization methods can solve monotone inclusion problems with complexity O(ε^{-2/(3p-1)}), which is provably optimal and improves prior bounds by using anchored extrapolation with Taylor approximations of operators.

This paper develops optimal higher-order methods for solving monotone inclusion problems—a fundamental class of optimization problems. The authors introduce the Anchored Extra-Proximal framework that achieves better convergence rates than prior methods by combining extrapolation with proximal updates. They prove their approach is optimal up to logarithmic factors across all orders.

training

Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers

Sep 24, 2026

Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek et al.

Training latent dynamics models for long-horizon stability requires explicitly optimizing for multi-step rollout accuracy, not just reconstruction—this restructures the solution space in ways that conventional metrics don't capture.

This paper shows that neural surrogate models for physics simulations fail during long predictions not because of poor compression, but because they're trained only to reconstruct data. The authors introduce training techniques—including Koopman operator learning and noise injection—that restructure the latent space to support stable long-horizon forecasting.

efficiencytrainingreasoning

Intrinsic-Extrinsic Coupling in Learning Dynamics

Sep 24, 2026

Qinyou Wang

A model's internal state and external training conditions interact in complex, non-additive ways—the same intervention can help or hurt depending on what happens next, which matters for understanding continual learning and model adaptation.

This paper studies how a machine learning model's current state interacts with future training dynamics. The authors develop methods to measure and manipulate learning states—like classifier weights and historical information—to understand when interventions help or hurt performance.

trainingevaluation

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

Sep 24, 2026

Nayoung Choi, Shengjian Chen, Xiaokai Wei et al.

Optimizing search query understanding components individually with search-engine-derived rewards outperforms single end-to-end optimization, showing that understanding how each component affects downstream retrieval matters more than just matching labels.

This paper presents a reinforcement learning framework for query understanding in search systems that optimizes multiple components (like intent classification and query expansion) separately using rewards from live search engine interactions, rather than treating it as a single end-to-end problem.

trainingreasoningapplications

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

Sep 24, 2026

Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh

Token-level tolerance to transcription ambiguity in ASR training reduces word error rates by ~9.5% on average by letting models skip disputed individual characters while keeping supervision for the rest of the word.

This paper addresses a real problem in speech recognition: reference transcripts often contain ambiguous pronunciations or spellings that the audio doesn't uniquely determine.

trainingevaluation

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Sep 22, 2026

Laizhen Li, Jiarui Li, Juanjuan Zhao et al.

You can move repetitive agent control logic from expensive LLM context into persistent, reusable code—cutting inference costs by 74-99% while keeping smaller models effective on complex tasks.

This paper introduces Growing Harness, a method that automatically builds reusable agent control code from task feedback instead of asking language models to repeatedly solve the same control problems. By learning executable code that handles recurring decisions, the approach reduces LLM calls by 76-92% while maintaining or improving task success rates across different model sizes.

agentsefficiencytraining

FleXray: Universal Clinical X-ray Segmentation

Sep 22, 2026

Victor Ion Butoi, Vivek Gopalakrishnan, John V. Guttag et al.

You can train powerful medical imaging models without expensive manual annotation by simulating realistic training data from existing 3D datasets—FleXray segments full-body X-rays and works on real clinical data despite being trained entirely on synthetic images.

FleXray is a generalist AI model that segments 60 anatomical structures in clinical X-rays across the entire body. Rather than manually labeling thousands of X-rays, the researchers built a physics-based simulator that generates realistic synthetic X-rays from existing 3D CT scans, then trained the model on these simulations.

trainingdataapplications

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Sep 22, 2026

Yuanteng Chen, Zhilei Liu, Peisong Wang et al.

On-policy distillation recovers reasoning capabilities in ultra-low-bit quantized models by training on the model's own generated outputs rather than fixed data, fixing the exposure bias problem that causes long-form reasoning to fail.

This paper tackles a critical problem in quantized language models: when you compress models to very low precision (under 3 bits), they lose the ability to do math and coding tasks because errors compound during long generation.

trainingefficiencyreasoning

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

Sep 21, 2026

Lei Yang, Mengyin Liu, Jia Wang et al.

Token-level correction during annotation is significantly faster than full rewriting and produces on-policy training data that preserves the model's natural generation patterns while providing precise supervision signals.

onPanda is an interactive annotation tool that helps create training data for AI models by letting annotators correct responses token-by-token. Instead of rewriting entire outputs, annotators find the first mistake, fix it, and let the model regenerate from that point.

trainingdataalignment

Harness-Zero: Harness Distillation via Agent-as-Harness

Sep 21, 2026

Haoran Ye, Yuxing Lu, Haonan Dong et al.

You can distill specialized harness behaviors into model weights by having an intermediate agent translate between different harness action spaces during training, letting you deploy with simpler harnesses while keeping performance gains.

This paper tackles how to transfer the benefits of specialized agent harnesses (external systems that improve model-environment interaction) into model weights so they work with simpler harnesses at deployment.

trainingagentsapplications

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Sep 21, 2026

Peng Xia, Rujun Han, Zifeng Wang et al.

Automatically improving agent harnesses through constrained evolution can boost performance while staying generalizable—the key is regularizing the search process to favor reusable mechanisms over task-specific tricks.

This paper presents RRSI, a method for automatically improving LLM agent systems by evolving their harnesses (prompts, tools, memory, control flow) while avoiding overfitting to training tasks. It uses regularization techniques like edit budgets and change filtering to find improvements that generalize to new benchmarks, achieving strong gains on both in-distribution and out-of-distribution tasks.

agentstrainingefficiency

Learning Physics from an Imperfect Ancestor

Sep 21, 2026

S. Mohammad Mousavi, Teeratorn Kadeethum, Nikolaos Bouklas et al.

Neural operators don't need to be accurate to be useful—they can guide PINNs away from spurious solutions by providing the right structural prior, enabling reliable PDE solving in regimes where either method alone would fail.

This paper shows how to combine neural operators (fast but inaccurate) with physics-informed neural networks (accurate but optimization-fragile) to solve PDEs reliably. An imperfect neural operator provides a structural hint about which solution the PINN should find, while the PDE residual refines it to high accuracy.

trainingreasoning
trainingevaluationapplications

CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

Sep 18, 2026

Bowen Ye, Lei Li, Shicheng Li et al.

Using source code itself as the primary input, you can automatically generate thousands of high-quality RL training tasks for coding agents without relying on manual annotations or development artifacts like issues.

CodeMidas automatically creates reinforcement learning training tasks from open-source code by using AI agents to explore codebases, generate test cases, and validate tasks. This approach scales coding agent training to 5,545 diverse tasks across 23 languages, improving performance on code repair, program synthesis, and terminal tasks by 8-18%.

trainingagentsreasoning

Benchmarking World Models for Continual Learning on Compositional Tasks

Sep 18, 2026

Haoyu Zhou, Joe Watson, Anson Lei et al.

Modular world model architectures better balance knowledge reuse with avoiding catastrophic forgetting in continual learning, but the field still lacks methods that effectively retain and reuse knowledge across sequential robot tasks.

This paper creates a benchmark to test how well world models (AI systems that learn to predict environment dynamics) can learn continuously across robot tasks without forgetting previous knowledge. The key innovation is using compositional tasks—where new tasks combine elements from earlier ones—to isolate what knowledge gets reused versus forgotten.

trainingevaluationarchitecture

Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise

Sep 18, 2026

Fabricio Breve

Pre-processing noisy labels with a lightweight particle-based algorithm before GCN training significantly improves robustness to label corruption while being faster than other robust methods.

This paper proposes PCC+GCN, a method that cleans noisy labels in graph data before training a Graph Convolutional Network. It uses a particle-based algorithm to identify and fix mislabeled nodes, then trains GCN on the refined labels. The approach is faster and more accurate than existing robust GCN methods across multiple datasets with different types of label noise.

trainingdata

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

Sep 18, 2026

Richard Zhe Wang

Attention heads need both the ability to abstain from attending and to filter noise from values—their importance shifts with model scale, suggesting future architectures should support both primitives.

This paper identifies two missing capabilities in standard softmax attention: abstention (allowing heads to output nothing instead of always producing weighted combinations) and noise filtering (suppressing interference from mixed features).

architectureefficiencytraining

Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision

Sep 17, 2026

Nitish Dashora, Douglas Chen, Idan Shenfeld et al.

By distilling VLM-identified task-salient information into a learned latent token during training, robots can efficiently handle long-horizon tasks at deployment time without expensive in-the-loop reasoning.

This paper introduces workspace tokens, a lightweight memory system for robotic manipulation that learns which task-relevant information to remember during training using a vision-language model, then uses this compressed memory at deployment without needing expensive VLM queries.

efficiencyagentstraining

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Sep 17, 2026

Ji Xie, Dewei Zhou, Xinyu Huang et al.

By combining real image supervision with pure-color anchors and a shared hex-prompt interface, you can now control exact object colors in both image generation and editing tasks with professional-grade precision.

Paint-Anything enables precise color control in AI image generation and editing by letting users specify exact colors using hex codes (like #FF5733). The system learns to understand hex values through a training dataset of 500K images with color labels and pure-color reference images, achieving much better color accuracy than existing methods.

multimodaltrainingapplications

How Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates?

Sep 17, 2026

Pochinapeddi Sai Bhargav, Nithin Somasekharan, Rohit Sunil Kanchi et al.

Pretraining neural PDE surrogates provides significant data efficiency gains (2-3x fewer samples needed), but this benefit shrinks or reverses when the target task involves different physics modeling than the pretraining source.

This paper investigates how pretraining neural networks to simulate fluid dynamics (PDE surrogates) helps when switching to new airfoil designs or physics models. The authors show that pretraining benefits depend on three factors: how much target data you have, how diverse that data is, and whether the source and target use different physics models.

trainingefficiencyevaluation

Score Centering Stabilizes Off-policy Reinforcement Learning

Sep 17, 2026

Martin Marek, Max Ryabinin

Score centering is a lightweight, composable fix for training-inference mismatch in RL that works by correcting accumulated bias rather than trying to eliminate the mismatch entirely—making it practical for large language models.

This paper identifies drift—a persistent bias that accumulates during training—as the main cause of instability when reinforcement learning models behave differently during training versus deployment. The authors propose 'score centering,' a simple mathematical correction that stabilizes training without requiring expensive changes to the inference engine.

trainingefficiencyalignment

RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

Sep 17, 2026

Yan Yu, Zhengxi Lu, Yizhou Liu et al.

Self-retiring distillation lets student agents learn from teachers strategically: absorb skills early when helpful, then graduate to independent learning once they've internalized enough knowledge to improve further on their own.

This paper proposes RetireOPD, a training method for AI agents that combines reinforcement learning with knowledge distillation from a teacher model. The key innovation is 'adaptive retirement'—the student agent automatically stops learning from the teacher once it becomes reliable enough, then continues improving on its own.

trainingagentsreasoning

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

Sep 17, 2026

Damiano Da Col, Maximilian Igl, Peter Karkus et al.

Using a privileged teacher (trained on high-level inputs like maps and bounding boxes) to supervise a camera-based student during closed-loop fine-tuning is 1000× more sample-efficient than direct RL post-training for autonomous driving.

OPTED improves autonomous driving policies by using a privileged teacher trained with reinforcement learning to guide a camera-based student model during closed-loop fine-tuning. This approach avoids expensive direct RL training in simulation while keeping the policy close to human demonstrations, achieving 1.6-9.5× improvements in driving performance.

trainingefficiencyapplications

dQwen3.5: Hybrid-Attention Diffusion Language Models

Sep 17, 2026

Anton Xue, Litu Rout, Aditya Akella et al.

Hybrid architectures combining attention and RNNs are surprisingly efficient starting points for diffusion language models, reaching comparable performance with 2x faster training than full-attention models.

This paper shows that hybrid-attention language models (mixing attention and RNN layers) can be efficiently adapted into diffusion language models, which generate text in any order rather than left-to-right. The dQwen3.5 models reach the same training loss in half the tokens compared to full-attention baselines, while maintaining strong performance in parallel decoding.

architecturetrainingefficiency

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

Sep 17, 2026

Thomas Steinecker, Denis Trescher, Alexander Bienemann et al.

Zero-shot sim-to-real transfer for autonomous driving is achievable by training on a consistent semantic representation in simulation and applying the same representation to real sensor data, eliminating the need for manual policy adaptation.

MILER is a reinforcement learning framework for autonomous driving that bridges simulation and real-world deployment without manual tuning. It trains policies in simulation using a semantic bird's-eye-view representation, then transfers them to real vehicles by converting camera and LiDAR data into the same representation format.

trainingapplicationsefficiency

Don't Mask the Environment: Observation Supervision Changes How Agents Explore Under RL

Sep 17, 2026

Juzheng Zhang, Disha Makhija, Manoj Ghuhan Arivazhagan et al.

Supervising observation tokens during SFT alongside action tokens improves downstream RL performance by preventing the policy from specializing away from environment modeling, enabling better exploration without extra data or parameters.

This paper shows that training language model agents to predict both actions AND environment observations (not just actions) leads to better exploration during reinforcement learning. Even though deployed agents never generate observations, learning to predict them helps the policy understand action consequences, resulting in higher pass rates on coding tasks without adding computational overhead.

trainingreasoning

Stable Movement for Nondual Lipschitz Convex Optimization: Efficiency and Nearly Optimal Oracle Rates

Sep 17, 2026

David Martínez-Rubio, Cristóbal Guzmán

By connecting convex optimization to the problem of chasing nested convex sets, the authors achieve nearly optimal oracle complexity for Lipschitz convex functions—meaning their algorithms are as efficient as theoretically possible with minimal wasted queries.

This paper develops efficient algorithms for convex optimization that achieve nearly optimal performance rates. The key innovation is reducing the problem to tracking moving convex sets, where the algorithm either finds good solutions or makes strategic cuts that force progress. The approach works for general norm constraints and achieves rates matching theoretical lower bounds.

training

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Sep 17, 2026

Zimu Han, Yiming Zeng, Jiyao Zhang et al.

You can improve robot manipulation models through human feedback without robot execution by using a handheld interface to detect when the policy is uncertain and identifying which parts of demonstrations are most important for learning.

This paper presents HIL-UMI, a method for improving vision-language-action robot models without needing a physical robot during training. Instead of repeatedly running the robot to collect new data, humans demonstrate tasks using a handheld interface while the system queries the current policy and intelligently decides when to collect new examples based on policy uncertainty and task progress.

trainingagentsefficiency

UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising

Sep 17, 2026

Kun Yao, Yuhang Zhou, Yichi Zhang et al.

Multi-objective alignment in generative models works better when you give each objective its own specialized parameters and policies rather than trying to optimize them all with a single reward signal.

UniPolicy is a framework for search advertising that balances multiple competing business goals—relevance, clicks, and revenue—by training separate specialized policies within a shared model.

trainingapplicationsreasoning

What Does Privileged Information Add to On-Policy Self-Distillation?

Sep 17, 2026

XiuYu Zhang, Wei Chow, Junfeng Fang et al.

Privileged information in self-distillation matters less than the cross-mode transfer between direct-response and thinking-enabled inference; most gains come from distillation itself, not from what the reference reveals.

This paper investigates what privileged information (like worked solutions or reasoning traces) actually adds to on-policy self-distillation in language models.

trainingreasoning

Objective vs. Search: Decomposing What Makes a Good Tokeniser

Sep 16, 2026

Ahmetcan Yavuz, Clara Meister, Tiago Pimentel

When building tokenizers for language models, the search procedure (how you build the vocabulary) matters more than the optimization objective (what you're trying to optimize), at least for compression efficiency—but neither strongly predicts linguistic performance.

This paper investigates what makes tokenizers (algorithms that break text into pieces for language models) effective by separating two design choices: what they optimize for (compression vs. likelihood) and how they search (bottom-up vs. top-down).

trainingefficiency

A Zeroth-Order Paradigm for LLM Preference Alignment

Sep 16, 2026

Peter Chen, Xi Chen, Wotao Yin et al.

ComPO offers a gradient-free alternative to direct preference optimization that may better handle preference pairs with small margins, with both theoretical convergence guarantees and empirical improvements across major LLM families.

This paper introduces ComPO, a new method for aligning language models with human preferences that uses comparison oracles instead of directly optimizing preference losses. Unlike standard approaches, ComPO extracts directional signals from preference pairs without computing gradients, and includes both offline and online variants with theoretical guarantees.

alignmenttrainingevaluation

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

Sep 16, 2026

Hejia Geng, Zesen Huang, Haoyang Li et al.

Scientific code repositories contain structured domain knowledge that can be systematically converted into agent training data, enabling models to learn both specialized scientific skills and general capabilities through verified interaction trajectories.

ScienceIDE converts scientific code repositories into learning environments for AI agents by automating the extraction of executable tasks, verification criteria, and domain knowledge. The system trains specialized models (PhAI-IDE family) on verified scientific code interactions, demonstrating that learning from scientific software improves both code repair and general reasoning capabilities.

agentstrainingapplications

ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents

Sep 15, 2026

Shuhan Xue, Jianyuan Zhong, Ziyuan Nan et al.

This work demonstrates a practical approach to building AI agents that improve through real-world use—by coupling task strategy refinement with model training, ScienceBuddy shows how interactive AI can evolve alongside the research it supports rather than remaining static.

ScienceBuddy is an AI research assistant that improves itself through a two-level feedback loop: it refines how it approaches tasks (harness evolution) while simultaneously training its underlying model on researcher feedback. By working directly in researchers' workflows, it learns from real scientific work and continuously adapts to become more useful.

agentsreasoningtraining

LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs

Sep 15, 2026

Thanapat Trachu, Samuele Cornell, William Chen et al.

Layer-specific compression in neural codecs reduces computational overhead by letting different quantization layers compress at their own optimal rates, improving efficiency without sacrificing quality.

This paper introduces LACE, a neural audio codec that compresses speech at different rates for each quantization layer rather than using a single compression step. By allowing each layer to have its own segmentation boundaries, LACE reduces sequence length and computational cost while maintaining audio quality.

efficiencyarchitecturetraining

Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback

Sep 15, 2026

Haichen Hu, Yuheng Zhang, David Simchi-Levi

You can distill LLMs without amplifying teacher bias by coupling teacher calibration with student training—using only source-domain feedback to iteratively correct the teacher while updating the student, achieving provable convergence without target rewards.

This paper addresses a key problem in LLM distillation: when you train a smaller model to mimic a larger one, you also copy the teacher's mistakes and biases. The authors propose Coupled Calibration and Learning (CCL), which alternates between correcting the teacher's errors using source-domain feedback and training the student on target questions.

trainingalignment