ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2567 papers25 this month12 topics
AllTraining 46Reasoning 42Efficiency 33Evaluation 32Agents 27Architecture 18Multimodal 17Applications 16Data 9Safety 9scaling 5Alignment 2

Oct 5 – Oct 11(5)

Direct Intermediate Initialization for Tilted Diffusion Samplers

Oct 5, 2026

Gregory D. Bellchambers

Initializing diffusion samplers at intermediate timesteps using pulled-back clean-space posteriors and Gaussian bridges can dramatically improve sample quality, especially when the posterior has modes that are rare under the prior.

This paper improves diffusion-based posterior sampling by initializing the sampler at an intermediate step rather than starting from pure noise. The key insight is that Gaussian-tilted targets along the reverse process can be reformulated as weaker clean-space posteriors, with samples transported analytically via a Gaussian bridge.

efficiency

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Oct 5, 2026

Benhao Huang, Chufan Shi, Junlin Chen et al.

Looped models can be dramatically more efficient by recognizing that recurrent states converge to fixed points, enabling truncated backpropagation, KV cache sharing, and faster RL—with learned depth priors and orthogonal injection providing better supervision than existing approaches.

This paper optimizes looped language models—models that process information through multiple recurrent passes—by leveraging fixed points in recurrent states.

training

Sep 28 – Oct 4(40)

From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing

Oct 2, 2026

Kuangyu Ding, Gesualdo Scutari

Graph decomposition into tree blocks enables more efficient decentralized optimization by jointly designing subproblems and communication patterns, with convergence rates that explicitly depend on network topology and function properties.

This paper develops a new framework for distributed optimization over networks where agents minimize functions while only communicating with neighbors. Instead of traditional mixing-based approaches, the method decomposes the network graph into tree-structured blocks, with agents cooperatively solving subproblems via message passing.

trainingefficiency

LESSER: Post-Training Data Selection with Output-Layer Gradients

Oct 2, 2026

Lyuxin David Zhang, Eric Wong, Surbhi Goel et al.

You can select effective training data for LLMs using only output-layer gradients from forward passes, cutting computation costs by ~10× compared to full-gradient methods without sacrificing downstream performance.

This paper proposes LESSER, a method for selecting high-quality training data for large language models by using only output-layer gradients instead of full-parameter gradients. By leveraging cheaper forward passes rather than expensive backward passes, LESSER reduces computational cost by 9.7× for supervised fine-tuning while maintaining performance comparable to full-gradient selection methods.

Sep 21 – Sep 27(21)

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

Sep 25, 2026

Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe et al.

Training language models to predict their confidence in intermediate reasoning steps—using only self-supervised learning—makes them generate shorter reasoning traces at inference time without any explicit length penalties or early-stopping mechanisms.

This paper shows that reasoning models can generate shorter, more efficient reasoning traces by learning to predict their own confidence in answers—without explicitly optimizing for length.

trainingefficiencyreasoning

New LoRA Skills Should Read but Never Write

Sep 25, 2026

Zeyan Li, Panqi Yang, Qirong Guo et al.

When combining multiple LoRA adapters, the internal representation and directional coupling between them matters more than the adapter weights themselves—fixing these choices lets you add skills sequentially without degrading previous ones.

This paper solves the problem of combining multiple fine-tuned LoRA adapters into a single model without interference. The key insight is that LoRA updates have multiple equivalent forms, and the choice matters when combining adapters.

Sep 14 – Sep 20(24)

An Interpretable Memory Decision Controller for LLM Agents Based on Three-Signal Complementarity: Decoupling Confidence and Consistency

Sep 18, 2026

Yiming Zhang, Jinghong Zhang, Haoran Zhao et al.

When using RAG with LLMs, blindly trusting all retrieved memories causes hallucinations; a lightweight geometric decision layer can filter unreliable memories without any learned parameters, making RAG safer and more trustworthy.

This paper introduces Memory Decision Layer (MDL), a parameter-free controller that decides whether to trust retrieved memories in RAG systems. It uses three signals—relevance, reliability, and task risk—combined through geometric operations to detect conflicting memories and prevent hallucinations, reducing errors by 56% when memories contradict each other.

reasoningsafetyefficiency

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

Sep 18, 2026

Richard Zhe Wang

Attention heads need both the ability to abstain from attending and to filter noise from values—their importance shifts with model scale, suggesting future architectures should support both primitives.

This paper identifies two missing capabilities in standard softmax attention: abstention (allowing heads to output nothing instead of always producing weighted combinations) and noise filtering (suppressing interference from mixed features).

Sep 7 – Sep 13(9)

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

Sep 10, 2026

Atindra Jha, Margaret Li, Jure Leskovec et al.

MoE models are significantly more vulnerable to data repetition than dense models—a critical concern as training data becomes scarce. Regularization helps, but the fundamental mismatch between sparsity and repeated data suggests new architectural approaches are needed.

This paper investigates how Mixture-of-Experts (MoE) models—which use sparse, specialized sub-networks—overfit more severely than dense models when training data is repeated.

trainingefficiencyscaling

On the Regularization Landscape for the Linear Recommendation Models

Sep 10, 2026

Dong Li, Zhenming Liu, Ruoming Jin et al.

Popular recommendation models that seem different actually rely on just two types of regularization (nuclear-norm or Frobenius-norm), and you can build better models by combining their strengths.

This paper analyzes why different deep learning-based recommendation algorithms perform similarly despite using different techniques. The researchers discovered that top-performing linear models all use either nuclear-norm or Frobenius-norm regularization. They propose new solutions combining the benefits of both approaches: low-rank structure with closed-form solutions and better expressiveness.

Aug 31 – Sep 6(1)

RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments

Sep 4, 2026

Quoc H. Nguyen, Ali Lafzi, Abhijeet Phatak et al.

Gradient-level personalization in federated learning works reliably across transformer architectures where parameter-level methods fail, enabling privacy-preserving retail systems that adapt to regional differences while maintaining strong performance.

RegionFed solves a critical problem in retail search: different regions have different query patterns and product preferences, but standard federated learning produces one-size-fits-all models that perform poorly everywhere.

trainingefficiencyapplications
efficiency
architecture

UniSlider: Perceptually Uniform Sliders for Continuous Image Editing

Oct 5, 2026

David Serrano-Lozano, Duygu Ceylan, Yannick Hold-Geoffroy et al.

Separating the UI slider from the underlying strength parameter and remapping it based on perceptual distance creates intuitive, predictable image editing interfaces that users prefer.

UniSlider makes image editing sliders feel natural by ensuring perceptual change increases smoothly and predictably as you move the slider. Current methods produce uneven results—some slider positions cause no visible change while others transform the image abruptly.

efficiencyapplications

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Oct 5, 2026

Haozhen Zhang, Haodong Yue, Quanyu Long et al.

Instead of pre-processing all memory upfront, MemPilot learns to make runtime decisions about memory curation, letting developers trade off accuracy against computational cost and speed based on their needs.

MemPilot is a framework that helps LLM agents manage memory more efficiently by deciding when to retrieve pre-stored information versus when to process raw conversation history on-demand.

agentsefficiencytraining

Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Oct 5, 2026

Sahil Mahendrakar

Knowledge distillation can compress speech synthesis models by 10x with minimal quality loss by separating the text-to-features and features-to-audio tasks and training them independently against a frozen teacher.

Paradee is a tiny text-to-speech model created by distilling a larger 82M-parameter teacher into just 8M parameters while keeping the same voice quality. The authors use a two-stage training approach: first synthesizing training data with the teacher model, then training separate text and audio components before combining them.

efficiencytrainingarchitecture
trainingefficiencydata

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

Oct 2, 2026

Hui Chen, Xuan Qi, James Xu Zhao et al.

By separating strategy exploration from implementation and reusing prompt prefixes across evolution steps, you can achieve better optimization results while spending 50-100x less on LLM API calls.

FrugalEvo optimizes LLM-guided program evolution by pairing a powerful LLM that explores strategies with a cheaper LLM that implements them, while using cache-efficient prompting to reduce costs. It introduces Budget-Aware AUC to measure solution quality per dollar spent, achieving state-of-the-art results on optimization tasks at a fraction of the cost of competing methods.

efficiencyreasoningagents

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Oct 2, 2026

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi et al.

By identifying and selectively training on high-impact token decisions rather than full sequences, you can make diffusion language models learn more efficiently with less data.

This paper introduces Pivot-SD, a training method for masked diffusion language models that focuses on the most impactful decisions during text generation. Instead of training on entire sequences, it identifies 'pivot' tokens—commitments that significantly reduce uncertainty about remaining words—and trains only on those, using success/failure signals to guide learning.

trainingefficiencyreasoning

MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

Oct 2, 2026

Sean Culatana, Shang-En Huang, Kang Li

MRVQ enables one quantized index to serve all dimension-rate combinations by using truncatable residual quantization, reducing memory overhead by 17.8-22x compared to training separate indices, though with modest quality trade-offs.

This paper introduces MRVQ, a quantization method that compresses embeddings for vector search while supporting flexible trade-offs between embedding dimension and compression rate. A single index can be truncated in two ways—dropping quantization stages or embedding coordinates—to adapt to different memory and latency constraints without storing multiple separate indices.

efficiencyevaluation

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Oct 2, 2026

Vladislav Gromadskii, David Li, Samson Gourevitch et al.

You can fine-tune fast diffusion generators for reward optimization without expensive reference rollouts by replacing intractable KL penalties with inverse-distillation regularization that provably bounds divergence.

IDRF is a method for fine-tuning masked discrete diffusion models (which generate sequences iteratively by predicting multiple tokens at once) to maximize rewards while staying close to a reference model. Instead of computing intractable likelihood penalties, it uses a clever regularization trick called inverse-distillation that upper-bounds the KL divergence.

trainingefficiencyreasoning

Embedding Prediction Helps Image Generation

Oct 1, 2026

Sihan Xu, Ji Xie, Zilin Wang et al.

Dynamically predicting and updating conditioning embeddings at each generation step improves diffusion model efficiency and quality—you don't need to reuse the same embedding throughout the entire denoising process.

This paper proposes using predicted embeddings as dynamic conditioning signals in diffusion transformers instead of static embeddings. A separate transformer (NEPA) predicts image embeddings at each denoising step, allowing the conditioning to adapt to the current noise level. The approach achieves competitive image generation quality on ImageNet with significantly less training compute.

architectureefficiencytraining

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Oct 1, 2026

Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.

Representing 3D shapes as overlapping 2D slices rather than voxels dramatically reduces computational cost while improving topological correctness—a practical win for efficient 3D generation at scale.

SILSA is a 3D generation framework that uses compact sliding-window slice representations instead of expensive voxel tokens to generate high-resolution 3D shapes. By encoding cross-sections along three axes and adding topology supervision based on persistence diagrams, it achieves better structural quality while using 70% fewer tokens and running 58% faster than competing methods.

architectureefficiency

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Oct 1, 2026

Jichao Jiang, Cristian McGee, El Houcine Bergou et al.

TACO cuts optimizer memory from 27.7 GB to 0.16 GB on 13B models by using a sparse, low-precision update strategy based on column-wise signs—making full-parameter fine-tuning practical on consumer GPUs without sacrificing model quality.

TACO is a new optimizer for fine-tuning large language models that dramatically reduces memory usage by storing only tiny gradient components per column instead of full optimizer state. It achieves 174× memory reduction compared to AdamW while maintaining accuracy, enabling fine-tuning of 30-32B models on a single GPU.

efficiencytraining

FERPO: Forward Entropy-Regularized Policy Optimization

Oct 1, 2026

Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

Using forward-KL instead of reverse-KL for policy fitting encourages broader exploration of high-value actions and avoids the computational cost of differentiating critics, leading to faster and more sample-efficient learning.

FERPO is a reinforcement learning algorithm that improves policies by using critic values directly rather than differentiating through them. It derives optimal target actions using entropy regularization and fits the actor to these targets using forward-KL divergence, which encourages exploring multiple high-value action modes while keeping importance weights stable.

trainingefficiencyreasoning

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Oct 1, 2026

Cristian McGee, El Houcine Bergou, Aritra Dutta

ZFO decouples direction selection from step-size selection in LLM fine-tuning, using gradient information plus two function evaluations to adaptively choose step sizes that often outperform fixed-step methods without the cost of full line searches.

This paper proposes ZFO, a lightweight optimization framework for fine-tuning large language models that intelligently selects step sizes by combining first-order gradient information with minimal zeroth-order function evaluations.

trainingefficiency

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Oct 1, 2026

Zhengming Yu, Junkun Yuan, Haotian Yang et al.

By framing distribution matching as a classification problem with discriminators, DMAD eliminates the memory overhead of auxiliary models while maintaining or improving generation quality—enabling practical few-step visual generation.

DMAD improves fast image and video generation by training lightweight student models to match teacher distributions without needing an auxiliary model. It uses two discriminator heads to learn density ratios directly, making the process more efficient while achieving state-of-the-art quality in one-to-four-step generation across images and videos.

efficiencytrainingarchitecture

Decoding Looped Transformers Better for (Almost) Free

Oct 1, 2026

Weihao Liu, Huangjie Zheng, Tianrong Chen et al.

Looped Transformers naturally produce weak-to-strong prediction pairs across recurrent passes; contrasting them during decoding improves quality and enables halving compute with no training needed.

This paper introduces LoopCD, a training-free decoding method for looped Transformers that reuses intermediate predictions from earlier recurrent passes to guide token selection. By contrasting predictions from different loop depths, LoopCD improves reasoning accuracy (e.g., AIME scores from 61.88% to 73.33%) while cutting inference compute by 22-48% through fewer required loops.

efficiency

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Oct 1, 2026

Joohwan Ko, Tetiana Parshakova, Diana Cai et al.

Quasi-Newton methods—traditionally limited to convex optimization—can now scale to deep learning by using variational objectives to ensure positive curvature and GPU-friendly matrix operations, outperforming Adam on ill-conditioned problems.

SoftServe is a new quasi-Newton optimization method for training deep neural networks that handles the challenges of non-convex optimization and massive parameter counts.

trainingefficiency

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Oct 1, 2026

Siqi Zhu, Suozhi Huang, Kaixuan Zhang et al.

When combining multiple RL-trained teachers into one student, the averaging method and optimizer choice matter more than raw gradient differences—response length weighting and precision loss can swing task performance by 2-5 percentage points.

This paper investigates how multiple teacher models transfer knowledge to a student model during on-policy distillation. The researchers found that loss averaging implicitly weights responses, Adam's optimizer smooths gradient differences, and low-precision arithmetic (BF16) masks small weight updates—with these factors significantly affecting which tasks the student learns best.

trainingefficiency

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

Oct 1, 2026

Juan S. Santillana

Don't trust tool-use benchmarks for small models—use verbatim-reproduction checks and token-probability probes to verify genuine capability before claiming tool use works.

Small language models can appear to use tools correctly on standard benchmarks while actually just memorizing training examples. This paper shows how keyword-matching tests miss real failures, proposes cheap diagnostic checks to catch false positives, and demonstrates a targeted fix that repairs tool-use ability in a 1.1B parameter model using minimal compute.

evaluationefficiency

Local Support Learning

Oct 1, 2026

Assaf Ben-Kish, Akarsh Kumar, James Glass et al.

You can prevent large models from forgetting old skills during new training by using a learnable gate that only activates weight updates when the input matches the current training distribution—no need to store old data.

This paper addresses catastrophic forgetting in large language models by treating it as a geometric problem in weight space.

trainingefficiencyalignment

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

Oct 1, 2026

Zilin Du, Bowen Yang, Boyang Albert Li

When scaling data selection with neural networks, the standard loss function causes poor generalization—a new loss function (PVM) that matches predicted values pointwise solves this and transfers better across datasets and model scales.

This paper addresses data selection for training large language models by proposing TESS, a framework that uses a neural network to score and select training examples. Unlike existing meta-learning approaches that assign per-sample weights, TESS uses a novel loss function (Pointwise Value Matching) that avoids optimization instability and improves generalization to new datasets and model sizes.

trainingdataefficiency

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

Oct 1, 2026

Yinheng Li, Justin Wagle

Modern LLMs are naturally good at making structured decisions from predefined options without fine-tuning, but targeted fine-tuning helps weaker models and specific tasks like routing—without degrading their conversational abilities.

This paper shows that large language models can already make categorical decisions (choosing from predefined options) without generating text, using their built-in token probabilities. The authors present LLM2Jev, a method to extract these decisions directly and optionally fine-tune models to improve decision-making on specific tasks, while keeping the model's text generation abilities intact.

trainingefficiencyapplications

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

Oct 1, 2026

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

HAO enables stable reinforcement learning under FHE encryption by preventing polynomial approximation errors from accumulating—achieving 0% boundary violations versus 83.8% for unprotected baselines, making privacy-preserving RL practically viable.

This paper solves a critical problem in privacy-preserving reinforcement learning: when you encrypt data with Fully Homomorphic Encryption (FHE) for cloud computation, you must replace nonlinear operations with polynomial approximations, which causes training to diverge.

safetyefficiencytraining

Image Classifiers are Efficient Self-Supervised Video Representation Learners

Sep 30, 2026

Owais Iqbal, Sudipta Sarkar, Shyam Marjit et al.

You can efficiently learn video representations by repurposing pretrained image models with clever masking strategies, avoiding expensive 3D architectures and reconstruction overhead.

VideoMSN uses standard image Vision Transformers to learn video representations without 3D models or reconstruction. It treats videos as grids of frames and masks either spatial patches or temporal frames, then aligns the two views using a Siamese loss. This achieves top results on video benchmarks while needing 32-160x fewer training epochs than prior methods.

efficiencymultimodal

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Sep 30, 2026

Razan El Mais, Ali Chehab, Ibrahim Issa et al.

When fine-tuning LLMs with differential privacy, untying input/output embeddings outperforms the standard weight-tied design and enables 60% memory savings—suggesting privacy-preserving training requires rethinking standard model architectures.

This paper investigates weight tying (sharing parameters between input and output embeddings) in large language models trained with differential privacy. The authors find that untying embeddings actually improves performance under DP-SGD, achieving up to 4.74% accuracy gains, while also enabling more memory-efficient privacy techniques.

safetyefficiencytraining

Turbo Harness: Instance-Adaptive Harness Optimization

Sep 30, 2026

Tunyu Zhang, Hao Wang, Kai Xu et al.

Adapting execution harnesses to individual task instances—rather than using a single global harness—consistently improves agent performance, and this adaptation can be automated by learning from previous optimization runs.

This paper introduces Turbo Harness, a system that automatically customizes AI agent execution frameworks (harnesses) for individual tasks by learning from past optimization runs. Instead of using one fixed harness for all tasks, it generates task-specific modifications that improve agent performance across diverse domains like interactive tasks, coding, and long-horizon planning.

agentstrainingefficiency

Scaling Laws for Looped Mixture of Experts

Sep 30, 2026

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

Looped MoE models combine two orthogonal efficiency axes: recurrence increases computational depth while sparsity expands capacity, and their scaling laws enable principled design of efficient models that match much larger dense models at the same compute budget.

This paper develops scaling laws that jointly model looped transformers (which use recurrence for computational depth) and Mixture-of-Experts (which use sparsity for capacity).

scalingefficiencyarchitecture

Looped Diffusion Transformer

Sep 30, 2026

Yong Xien Chng, Tianyi Chen, Wenwen Tong et al.

Looping computation within denoising steps is more efficient than scaling model size—you can get better images with fewer parameters by iteratively refining representations through repeated block execution.

This paper proposes Looped Diffusion Transformer, which improves text-to-image generation by repeatedly processing the same Transformer blocks within each denoising step rather than making models larger.

efficiencyarchitecturescaling

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Sep 29, 2026

Bingchen Yao, Haobo Xu, Haokun Lin et al.

Quantization errors in recurrent states don't matter equally: errors in long-lived memory and in dimensions that strongly influence outputs cause more accuracy loss, so allocating precision based on these factors enables aggressive compression without sacrificing performance.

STEPQuant is a quantization method that compresses the persistent memory states used in linear attention models. By analyzing how quantization errors affect model outputs differently across time and space, the method allocates precision strategically—giving more bits to memory that lasts longer and to state dimensions that matter most for predictions.

efficiency

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Sep 29, 2026

Yi Pan, Haocheng Xi, Kan Zhu et al.

You can compress linear attention's recurrent state to 8-bit without significant quality loss by quantizing only at window boundaries and preserving outliers as special tokens—enabling faster inference on long sequences.

LeapQuant reduces the inference cost of linear attention models by quantizing their recurrent state to 8-bit precision while maintaining accuracy. It uses per-window quantization to limit error buildup and compensator tokens to handle outliers, achieving 2-3.7x speedups on real hardware.

efficiencytraining

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Sep 29, 2026

Kuan-Po Huang, Haohe Liu, Puyuan Peng et al.

Emotion vectors in TTS models can be decomposed into a neutral-shift component and an emotion-specific component—controlling them separately via steering achieves much better emotion control than treating them as a single direction.

This paper improves emotional speech generation by decomposing emotion vectors into shared and residual components, then controlling them separately without retraining the model. The method, EmoRES, significantly outperforms prior vector steering approaches on multiple emotion metrics and human evaluation.

efficiencymultimodaltraining

AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

Sep 29, 2026

Rishabh Agrawal, Hejie Cui, Shasha Li et al.

Selective learning from feedback—keeping corrections that significantly change model behavior while filtering out those that don't—improves advisor performance and generalization to new tasks and model families.

This paper presents AdviSD, a method for training small AI advisors that guide frozen large language models through natural-language feedback. The key innovation is selectively learning from corrections based on how much they actually change the executor's behavior, avoiding learning from corrections that don't meaningfully affect outcomes.

trainingreasoningefficiency

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Sep 29, 2026

Yu Xu, Yuxin Zhang, Xiao Yang et al.

Video models need different scaling strategies than language models—splitting experts by semantic role and allowing imbalanced routing produces better video generation than forcing uniform expert usage.

This paper proposes SplitMoE, a new way to scale video generation models using a split expert architecture that avoids forcing uniform expert usage.

scalingarchitectureefficiency

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

Sep 29, 2026

Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen et al.

Long-context evaluation needs to measure both accuracy and computational efficiency—the same model can be dramatically more or less efficient depending on its processing strategy, and current benchmarks don't capture this variation.

This paper introduces LongHarness Bench, a benchmark for evaluating how well language models handle long documents using different processing strategies. The benchmark tests both accuracy and efficiency, requiring models to find relevant information across scattered context and reason strategically rather than reading everything.

evaluationefficiency

Telescopic Language Models

Sep 28, 2026

Zhilin Guo, Boqiao Zhang, Hakan Aktas et al.

You can train one model that works well at any depth by supervising random layer prefixes during training—no architectural tricks needed, just a smarter training objective that makes the model elastic across compute budgets.

This paper introduces Telescopic Language Models (TLMs), which train a single model that works effectively at every layer depth rather than requiring separate models for different compute budgets.

trainingefficiencyarchitecture

PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

Sep 28, 2026

Zimo Wang, Junkun Yuan, Angtian Wang et al.

Critic error accumulation is a fundamental bottleneck in distilling video diffusion models; filtering it via projection dramatically improves sample quality without architectural changes or extra computation.

This paper improves video diffusion model distillation by fixing a key problem: critic errors that accumulate during training and degrade sample quality. PDMD uses a simple mathematical projection to filter out these errors while preserving useful learning signals, achieving better video quality with fewer computational steps—all with just a one-line code change.

efficiencytrainingevaluation

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Sep 28, 2026

Chaoqian Ouyang, Ling Yue, Libin Zheng et al.

Token consumption in agentic LLM workflows is unpredictable and can vary 10x+ per task—TokenCast forecasts it accurately by tracking execution segments and context growth, improving budget planning by 14.5% on average.

TokenCast predicts how many tokens an LLM agent will consume during task execution, which varies wildly across runs due to tool use and growing context. It learns cost patterns for each execution step and updates predictions as the agent runs, enabling better budget control without extra LLM calls.

agentsefficiencyevaluation

Neural Harmonic Measure Operator

Sep 28, 2026

Jinjin He, Sinan Wang, Yuchen Sun et al.

NHMO enables fast, geometry-aware PDE solving by learning a domain-specific kernel once, then reusing it for any boundary conditions or sources—eliminating expensive retraining for each new problem.

This paper presents Neural Harmonic Measure Operator (NHMO), a neural network approach that solves elliptic PDEs (like Laplace and Poisson equations) on domains with varying shapes.

reasoningefficiencyarchitecture

How to Loop MoE: Flatten the Experts, Untie the Attention

Sep 28, 2026

Shouren Wang, Chuang Ma, Mohsen Hariri et al.

To build better looped MoE models, flatten the expert hierarchy (more experts per layer, more passes) and give each pass independent attention—this lets tokens access more experts while maintaining computational efficiency.

This paper improves looped mixture-of-experts (MoE) models by flattening the expert structure and untying attention parameters. The key insight is that by doubling experts per layer and doubling passes through the network while keeping compute fixed, models can route tokens through more diverse experts, improving performance.

architectureefficiencytraining

KV-streams for Efficient Compaction in Agentic Reinforcement Learning

Sep 28, 2026

Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda et al.

KV-streams enables efficient scaling of agentic LLMs to longer horizons by streaming cached computations rather than recomputing them, making long-context RL training practical without sacrificing performance.

This paper introduces KV-streams, a technique that speeds up training of long-horizon agentic language models by streaming the key-value cache forward during context compaction instead of repeatedly refilling it. The method achieves 2.6-5x training speedup while maintaining performance, and shows that the streamed cache can retain information beyond the visible context window.

efficiencytrainingagents

Towards Communication-Efficient Social Intelligence in Language Agents

Sep 28, 2026

Linxiao Gong, Yijie Xu, Tianfu Wang et al.

TACT enables language agents to achieve better social outcomes while using fewer tokens and messages by having specialized teachers refine communication strategy and expression, then distilling improvements into the student agent.

This paper introduces Teacher-Assisted Communication Training (TACT), a method that helps language agents communicate more efficiently during social interactions. TACT improves how agents negotiate and coordinate by having a teacher refine both what agents say (expression) and how they say it (strategy), then distills these improvements back into the student agent.

trainingagentsefficiency

Improving Test-Time Scaling with Adaptive Looped Transformers

Sep 28, 2026

Yichen You, Tianyu Fu, Aosong Feng et al.

Adaptive token-level iteration selection during inference can significantly improve test-time scaling efficiency—TaH2 achieves 53% better accuracy-per-compute gains than fixed looping by intelligently deciding which tokens deserve extra processing passes.

This paper improves how AI models use extra computation time during inference by introducing TaH2, which selectively applies multiple processing passes to tokens that benefit most from them. Unlike standard looped transformers that process every token multiple times, TaH2 learns which tokens need extra iterations, achieving better accuracy gains per unit of compute on math reasoning tasks.

efficiencyreasoning

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Sep 28, 2026

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan et al.

When adapting pre-trained vision models to use linear attention, directly copy MLP weights but distill attention behavior—this simple strategy closes the performance gap between efficient and standard transformers.

This paper shows how to initialize linear Vision Transformers (efficient attention models) using weights from standard Softmax ViTs. The key insight: copy the MLP layers directly since they learn general representations, but use distillation to transfer the attention mechanism since it's operator-specific.

efficiencyarchitecturetraining
trainingefficiency

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Sep 25, 2026

Varun Reddy, Bernardo L. Sabatini, Houman Safaai

Common-mode error in direct feedback alignment causes training stalls by saturating hidden units; this can be prevented by centering batch errors or calibrating the readout baseline, enabling faster learning without changing the core algorithm.

Direct feedback alignment trains neural networks using fixed random error projections, but gets stuck learning near a baseline predictor. The paper identifies that a shared error component across inputs drives hidden units to saturation, slowing learning. Simple fixes like centering errors or adjusting the baseline readout can prevent this collapse and speed up training.

trainingefficiency

DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

Sep 25, 2026

Quang Nguyen, Hieu Nguyen, Hien Hoang et al.

Building effective AI tutors for non-English regions requires both technical efficiency (faster inference on consumer hardware) and domain-specific learning (accumulating local knowledge from real interactions rather than relying on pre-training).

DeepEdu-v1 is an AI tutoring system for Vietnamese students that runs locally to protect data privacy and avoid hallucinations from Western-trained models. It uses two key innovations: a smarter way to handle long conversations that reduces processing time by 35%, and a self-improving system that learns from past tutoring interactions instead of requiring expensive retraining.

efficiencyapplicationsagents

EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

Sep 25, 2026

Kunxiong Zhu, Zhihao Shu, Hangyu Zheng et al.

Multimodal LLM serving requires rethinking GPU resource allocation around the Encode stage—treating it as a bottleneck control point rather than a separate service unlocks significant throughput gains.

EAServe optimizes serving multimodal LLMs by treating the Encode stage as a control point for the three-stage Encode-Prefill-Decode pipeline. It uses adaptive micro-batching, partial offloading, and GPU partitioning to balance resource utilization across stages, achieving 4.3x higher throughput than existing systems under latency constraints.

efficiencymultimodalarchitecture

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

Sep 24, 2026

Jiabin Qiu, Zixuan Chen, Hongye Cao et al.

World models for planning need to preserve action-discriminative information through training, not just minimize prediction error—this simple insight significantly improves both simulation and real-world robotic control performance.

This paper shows that world models trained only to predict what actually happens often fail at model predictive control, which requires comparing different action choices.

reasoningagentsefficiency

Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning

Sep 24, 2026

Sudip Bhujel, Shanghao Shi, Ruiquan Huang et al.

Distributed RL agents that share only gradients—not raw data—still leak sensitive trajectory information through temporal correlations; defending against this requires sequence-aware privacy mechanisms, not just per-step protections.

This paper reveals a critical privacy vulnerability in distributed embodied AI systems. When agents send policy gradients to a server instead of raw sensor data, attackers can reconstruct the agent's complete trajectory of observations and actions by analyzing the temporal patterns in these gradients.

safetyagentsefficiency

Rolling-WAM: World Action Models with Rolling Imagination

Sep 24, 2026

Yinghua Zhou, Junjie Ye, Yiqi Zhao et al.

Distributing prediction computation across replanning cycles via a rolling noise schedule achieves 4.5x speedup in robot control latency without sacrificing task performance.

Rolling-WAM speeds up robot control by spreading the computation of predicting future actions and images across multiple planning cycles instead of doing it all at once. Instead of fully planning the entire future from scratch each time, it maintains a sliding window of partially-computed predictions at different stages, letting them gradually refine as new camera data arrives.

efficiencyagentsreasoning

PoEM: Predicting RL Outcomes from Existing Policies

Sep 24, 2026

Kimia Hamidieh, Giannis Daras, Antonio Torralba

You can predict RL outcomes for new reward functions by combining existing trained models mathematically, avoiding the computational cost of retraining—useful when experimenting with different objectives or combining multiple goals.

PoEM predicts what a reinforcement learning model will do with a new reward function by combining existing models trained on different rewards, without running expensive RL training. The method works by finding that RL policies live in a low-rank space that can be reconstructed as a linear combination of existing policies.

trainingefficiencyreasoning

Minimally Invasive Steering of Language Models

Sep 24, 2026

Taha Entesari, Jingyu Zhang, Daniel Khashabi et al.

You can adapt a frozen language model to new rewards at inference time by carefully controlling how much you perturb its hidden states—using Fisher information to measure and limit distributional changes prevents quality degradation.

This paper introduces MISVO, a technique for steering frozen language models at test time by adding vectors to hidden states while minimizing unwanted changes to output quality. Using Fisher information geometry, the method penalizes interventions that distort the token distribution, enabling efficient reward optimization without retraining the model.

efficiencyalignment

Beyond Compression: Training Latent Representations for Stable Long-Horizon Rollout in Neural Surrogate Solvers

Sep 24, 2026

Andreas E. Robertson, Ashley T. Lenau, John D. Shimanek et al.

Training latent dynamics models for long-horizon stability requires explicitly optimizing for multi-step rollout accuracy, not just reconstruction—this restructures the solution space in ways that conventional metrics don't capture.

This paper shows that neural surrogate models for physics simulations fail during long predictions not because of poor compression, but because they're trained only to reconstruct data. The authors introduce training techniques—including Koopman operator learning and noise injection—that restructure the latent space to support stable long-horizon forecasting.

efficiencytrainingreasoning

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

Sep 24, 2026

Linghua Zhang

Decoupling VLM planning from action execution using a lightweight executor reduces model serving costs dramatically without sacrificing task performance on mobile GUI automation.

Jev-Mobile improves mobile GUI agents by separating planning from execution: a vision-language model makes high-level decisions infrequently, while a lightweight decision model (Jev) handles repeated low-level actions. This cuts inference costs by 73% and execution time by 33% while maintaining 79% task success on Android tasks.

agentsefficiencyapplications

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Sep 22, 2026

Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen

Diffusion LLMs can be dramatically accelerated by jointly optimizing memory I/O patterns in KV caching and using the model itself for draft-and-verify decoding, rather than treating these optimizations separately.

This paper presents Flash-dLLM, a system that speeds up diffusion language models (an alternative to traditional autoregressive LLMs) by optimizing how the model stores and reuses computation results during inference.

efficiencyarchitecture

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Sep 22, 2026

Trang Nguyen, Eulrang Cho, Bingqing Chen et al.

By compacting context through selective truncation rather than rephrasing, agents can maintain performance on long-horizon coding tasks while reducing inference costs significantly—making test-time scaling more economical.

CliffCompaction is a technique that compresses long conversation histories for AI coding agents by selectively removing less important content while keeping everything else unchanged. This cuts costs by up to 50% while maintaining performance, enabling agents to work on complex coding problems that span millions of tokens across multiple sessions.

efficiencyagentsreasoning

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Sep 22, 2026

Jennifer Williams, Dave Farris, Jeff Farris et al.

AI agents can complete inference engineering tasks locally, but production correctness is much harder—end-to-end serving tests catch failures that other tests miss, exposing a critical gap between development and deployment.

SWE-Serve is a benchmark with 53 real production tasks from SGLang that tests whether AI agents can implement inference serving features correctly—not just locally, but in production. It reveals a major gap: one-third of code changes that pass unit tests fail end-to-end serving tests, showing that current agents struggle with production-grade correctness.

evaluationagentsefficiency

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Sep 22, 2026

Laizhen Li, Jiarui Li, Juanjuan Zhao et al.

You can move repetitive agent control logic from expensive LLM context into persistent, reusable code—cutting inference costs by 74-99% while keeping smaller models effective on complex tasks.

This paper introduces Growing Harness, a method that automatically builds reusable agent control code from task feedback instead of asking language models to repeatedly solve the same control problems. By learning executable code that handles recurring decisions, the approach reduces LLM calls by 76-92% while maintaining or improving task success rates across different model sizes.

agentsefficiencytraining

Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows

Sep 22, 2026

Remy Stewart, Olabode Anise, Andrew Hogan et al.

AI design tools deliver real time savings (~20%) in controlled settings, but benefits vary by user expertise—product managers gain more than professional designers, indicating task-dependent value.

Researchers tested whether AI-powered design tools (Figma Make) actually save time by having 50 designers and 50 product managers complete design tasks with and without the tool. They found about 20% faster completion times with AI assistance, especially for non-designers, suggesting these tools may let product managers do more design work themselves.

evaluationapplicationsefficiency

Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning

Sep 22, 2026

Yuanteng Chen, Zhilei Liu, Peisong Wang et al.

On-policy distillation recovers reasoning capabilities in ultra-low-bit quantized models by training on the model's own generated outputs rather than fixed data, fixing the exposure bias problem that causes long-form reasoning to fail.

This paper tackles a critical problem in quantized language models: when you compress models to very low precision (under 3 bits), they lose the ability to do math and coding tasks because errors compound during long generation.

trainingefficiencyreasoning

LoRA-generating hypernetworks for efficient on-device LLM generative personalization

Sep 21, 2026

Sean Augenstein, Li Ding, Jihwan Lee et al.

Hypernetworks can efficiently generate personalized LoRA adapters on mobile devices by mapping user context to weight modifications, avoiding the latency costs of context extension while remaining computationally feasible for resource-constrained devices.

This paper presents a method for personalizing language models on mobile devices by training a hypernetwork that generates customized LoRA (low-rank adaptation) weights based on a user's context.

efficiencyapplications

RRSI: Regularized Recursive Self-Improvement of Agent Harnesses

Sep 21, 2026

Peng Xia, Rujun Han, Zifeng Wang et al.

Automatically improving agent harnesses through constrained evolution can boost performance while staying generalizable—the key is regularizing the search process to favor reusable mechanisms over task-specific tricks.

This paper presents RRSI, a method for automatically improving LLM agent systems by evolving their harnesses (prompts, tools, memory, control flow) while avoiding overfitting to training tasks. It uses regularization techniques like edit budgets and change filtering to find improvements that generalize to new benchmarks, achieving strong gains on both in-distribution and out-of-distribution tasks.

agentstrainingefficiency

Rare Event Estimation via Iterative Unalignment

Sep 21, 2026

Hanming Yang, Daksh Mittal, Jing Dong et al.

For safe AI deployment, you need to estimate how often catastrophic rare events occur in agent behavior. This paper provides a practical method using importance sampling with learned weight perturbations, achieving massive efficiency gains over naive approaches.

This paper tackles estimating extremely rare event probabilities in AI agent trajectories—events so uncommon that standard Monte Carlo sampling is impractical. The authors develop a new importance sampling method that tweaks a language model's weights to generate more likely rare events, using gradient-based optimization to search efficiently.

safetyevaluationefficiency
architectureefficiencytraining

Workspace Models: Lightweight Robotic Memory via Saliency-Driven Supervision

Sep 17, 2026

Nitish Dashora, Douglas Chen, Idan Shenfeld et al.

By distilling VLM-identified task-salient information into a learned latent token during training, robots can efficiently handle long-horizon tasks at deployment time without expensive in-the-loop reasoning.

This paper introduces workspace tokens, a lightweight memory system for robotic manipulation that learns which task-relevant information to remember during training using a vision-language model, then uses this compressed memory at deployment without needing expensive VLM queries.

efficiencyagentstraining

How Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates?

Sep 17, 2026

Pochinapeddi Sai Bhargav, Nithin Somasekharan, Rohit Sunil Kanchi et al.

Pretraining neural PDE surrogates provides significant data efficiency gains (2-3x fewer samples needed), but this benefit shrinks or reverses when the target task involves different physics modeling than the pretraining source.

This paper investigates how pretraining neural networks to simulate fluid dynamics (PDE surrogates) helps when switching to new airfoil designs or physics models. The authors show that pretraining benefits depend on three factors: how much target data you have, how diverse that data is, and whether the source and target use different physics models.

trainingefficiencyevaluation

Score Centering Stabilizes Off-policy Reinforcement Learning

Sep 17, 2026

Martin Marek, Max Ryabinin

Score centering is a lightweight, composable fix for training-inference mismatch in RL that works by correcting accumulated bias rather than trying to eliminate the mismatch entirely—making it practical for large language models.

This paper identifies drift—a persistent bias that accumulates during training—as the main cause of instability when reinforcement learning models behave differently during training versus deployment. The authors propose 'score centering,' a simple mathematical correction that stabilizes training without requiring expensive changes to the inference engine.

trainingefficiencyalignment

GeoAAC: Geometry-Based Adaptive Action Chunking from Denoising Trajectories in VLA Policies

Sep 17, 2026

Xin Chen, Sen Chen, Yujuan Ding et al.

By analyzing the geometric properties of diffusion-based trajectory generation, you can detect prediction uncertainty and dynamically adjust action planning horizons without retraining—improving robotic control performance by up to 8.7 percentage points.

This paper introduces GeoAAC, a method that dynamically adjusts how many steps ahead a robot should plan based on task difficulty. Instead of using a fixed planning horizon, it analyzes the geometry of the prediction process to detect when the model is uncertain, then shortens or lengthens the planning window accordingly.

reasoningefficiencyagents

Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

Sep 17, 2026

Hanchu Zhou, Brendan Lynch, Raman Goyal et al.

By treating vision and tactile signals at different timescales in a lightweight architecture, you can build fast tactile-aware robot controllers (11.9ms latency) that outperform larger models on real contact-rich tasks.

This paper presents Agile-WAM, a tactile world action model that predicts future robot states and actions for contact-rich manipulation tasks.

architecturemultimodalefficiency

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

Sep 17, 2026

Damiano Da Col, Maximilian Igl, Peter Karkus et al.

Using a privileged teacher (trained on high-level inputs like maps and bounding boxes) to supervise a camera-based student during closed-loop fine-tuning is 1000× more sample-efficient than direct RL post-training for autonomous driving.

OPTED improves autonomous driving policies by using a privileged teacher trained with reinforcement learning to guide a camera-based student model during closed-loop fine-tuning. This approach avoids expensive direct RL training in simulation while keeping the policy close to human demonstrations, achieving 1.6-9.5× improvements in driving performance.

trainingefficiencyapplications

dQwen3.5: Hybrid-Attention Diffusion Language Models

Sep 17, 2026

Anton Xue, Litu Rout, Aditya Akella et al.

Hybrid architectures combining attention and RNNs are surprisingly efficient starting points for diffusion language models, reaching comparable performance with 2x faster training than full-attention models.

This paper shows that hybrid-attention language models (mixing attention and RNN layers) can be efficiently adapted into diffusion language models, which generate text in any order rather than left-to-right. The dQwen3.5 models reach the same training loss in half the tokens compared to full-attention baselines, while maintaining strong performance in parallel decoding.

architecturetrainingefficiency

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

Sep 17, 2026

Thomas Steinecker, Denis Trescher, Alexander Bienemann et al.

Zero-shot sim-to-real transfer for autonomous driving is achievable by training on a consistent semantic representation in simulation and applying the same representation to real sensor data, eliminating the need for manual policy adaptation.

MILER is a reinforcement learning framework for autonomous driving that bridges simulation and real-world deployment without manual tuning. It trains policies in simulation using a semantic bird's-eye-view representation, then transfers them to real vehicles by converting camera and LiDAR data into the same representation format.

trainingapplicationsefficiency

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Sep 17, 2026

Haocheng Xi, Yiming Xie, Hexu Zhao et al.

Hybrid attention architectures that combine local Softmax with linear memory can dramatically accelerate video generation without sacrificing quality—enabling practical real-time video synthesis on modern hardware.

Video DeltaNet combines local attention with efficient linear memory to speed up video diffusion models. By mixing Softmax attention for fine details with a new Video Delta Attention mechanism for long-range context, it generates high-quality videos 14.5x faster than baseline models while maintaining visual quality.

efficiencyarchitecturemultimodal

On-Demand Attention: Language Models Know When to Recall

Sep 17, 2026

Haibo Feng, Ruiqi Liang, Hanyang Peng et al.

Pretrained models already know which tokens matter for the next prediction—you can train a lightweight module to detect this and skip expensive full-context reads, cutting inference cost without retraining the base model.

This paper shows that language models can predict when they need to read their full context history during generation. The authors introduce On-Demand Attention (ODA), which uses a small trained module to decide when to use expensive global attention versus cheaper local attention, reducing computation while maintaining quality on long documents.

efficiency

Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models

Sep 17, 2026

Frank E. Bobe, Gregory D. Vetaw, Darshan W. Bryner et al.

Activation steering can be automated to find optimal intervention points in transformers, but practitioners should be aware that stronger steering increases vulnerability to prompt injection attacks—a critical concern for deployed agent systems.

Deep Noir automatically discovers where and how to steer LLM activations to change model behavior without retraining. Using a technique called Logit Lens to find optimal intervention points, the system improves spam detection by up to 42 percentage points across different model sizes.

safetyefficiency

RISC-V and machine learning: a survey

Sep 17, 2026

Shriman Keshri, Apparna Singh, Chinmaya Kumar Palo et al.

RISC-V offers a customizable, open alternative to proprietary chips for ML, but realizing its potential requires standardized extensions, better toolchain maturity, and specialized accelerator designs tailored to neural network workloads.

This survey examines how RISC-V, an open-source processor architecture, is being adapted for machine learning workloads. It analyzes existing implementations, software tools, and real-world applications, identifying strengths in energy efficiency and instruction extensions while highlighting challenges like fragmentation and standardization gaps.

architectureefficiency

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Sep 17, 2026

Zimu Han, Yiming Zeng, Jiyao Zhang et al.

You can improve robot manipulation models through human feedback without robot execution by using a handheld interface to detect when the policy is uncertain and identifying which parts of demonstrations are most important for learning.

This paper presents HIL-UMI, a method for improving vision-language-action robot models without needing a physical robot during training. Instead of repeatedly running the robot to collect new data, humans demonstrate tasks using a handheld interface while the system queries the current policy and intelligently decides when to collect new examples based on policy uncertainty and task progress.

trainingagentsefficiency

COIN-GP: Cooperative Online Learning in Networked Distributed Systems with Partial Measurements via Gaussian Process Regression

Sep 17, 2026

Zewen Yang, Xiaobing Dai, Zhenxiao Yin et al.

When building distributed sensor networks with incomplete data, combining state observers with cooperative Gaussian Process learning can accurately estimate both what's happening in the system and how it behaves, with proven error bounds.

This paper presents a method for distributed systems with multiple sensors to estimate both system states and unknown dynamics when only partial measurements are available. It combines observer-based estimation with online Gaussian Process regression across a network, includes a smart data collection strategy, and provides theoretical guarantees on estimation accuracy.

reasoningefficiency

Objective vs. Search: Decomposing What Makes a Good Tokeniser

Sep 16, 2026

Ahmetcan Yavuz, Clara Meister, Tiago Pimentel

When building tokenizers for language models, the search procedure (how you build the vocabulary) matters more than the optimization objective (what you're trying to optimize), at least for compression efficiency—but neither strongly predicts linguistic performance.

This paper investigates what makes tokenizers (algorithms that break text into pieces for language models) effective by separating two design choices: what they optimize for (compression vs. likelihood) and how they search (bottom-up vs. top-down).

trainingefficiency

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Sep 16, 2026

Zixi Chen, Akshay Vegesna, Samip Dahal et al.

Architectural design choices like recursive depth and boundary operators can fundamentally change scaling laws, enabling exponential compute efficiency improvements—challenging the idea that scaling laws are fixed.

This paper shows that architectural changes—specifically model growth through recursive depth (looping) and boundary operators—can improve how efficiently transformers use compute during training. A 7.4B looping model matches GPT-3 13B's performance with 20× less computation, with efficiency gains that grow larger at bigger scales.

architecturescalingefficiency

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

Sep 16, 2026

Kaijun Zhou, Zhiyang Li, Le Chen et al.

Robots doing repetitive tasks can reuse cached visual features and neuron activations from previous executions, cutting VLA inference time by 30-40% without sacrificing accuracy—critical for real-time robot responsiveness.

This paper presents rMuscle, a framework that speeds up Vision-Language-Action (VLA) models for robot control by caching and reusing computation across repeated tasks.

efficiencyapplicationsagents

What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity

Sep 15, 2026

Congjing Zhang, Vashishtha Patil, Henning Lange et al.

When deploying pruned LLMs for tool use, mixture-of-experts architectures are significantly more robust than dense models, and you need to evaluate specific action components—not just overall accuracy—to catch degradation early.

This paper studies how pruning (removing parts of neural networks to reduce size) affects large language models' ability to control smart home devices. The researchers test four different LLMs with various pruning strategies, finding that dense models break suddenly after modest pruning, while mixture-of-experts models are more robust.

efficiencyevaluationapplications

LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs

Sep 15, 2026

Thanapat Trachu, Samuele Cornell, William Chen et al.

Layer-specific compression in neural codecs reduces computational overhead by letting different quantization layers compress at their own optimal rates, improving efficiency without sacrificing quality.

This paper introduces LACE, a neural audio codec that compresses speech at different rates for each quantization layer rather than using a single compression step. By allowing each layer to have its own segmentation boundaries, LACE reduces sequence length and computational cost while maintaining audio quality.

efficiencyarchitecturetraining

JustFit: 200K-Token LLM Serving on a 24 GiB Laptop with Just-in-Time State Management

Sep 15, 2026

Yuhua Chen

Smart memory management—not just model compression—can unlock 200K-token context on consumer hardware, making local AI development practical for reasoning tasks.

JustFit enables running large language models with 200K-token context on a 24GB laptop by intelligently managing memory. It uses three techniques: compressing key-value cache data, swapping model components between storage and RAM, and preserving state across requests. Tests show it can handle 6.93x more context than existing methods while maintaining reasonable speed.

efficiency

Decomposition Buys Integrity, Not Yield

Sep 15, 2026

Rong He

Multi-agent decomposition inherently loses information as it scales—each tier passes up only a fraction of findings—but this cost is sometimes worth paying for reduced context and compute, especially when the root's memory becomes the bottleneck.

This paper analyzes how splitting tasks across multiple agents in a tree structure affects information flow from leaf agents back to the root. The authors model this as a probabilistic process and find that decomposition trades yield (fewer findings reach the top) for benefits like reduced context size and lower computational cost.

agentsreasoningefficiency

Reduced-Space Multi-Fidelity Bayesian Optimization of Process Simulation Models

Sep 15, 2026

Niki Triantafyllou, Andrea Bernardi, Maria M. Papathanasiou

By combining sensitivity analysis for dimension reduction with multi-fidelity optimization, you can reduce expensive simulator calls by 50%+ while maintaining optimization quality—critical for industrial design where each simulation costs hours or days.

This paper presents a method to optimize complex industrial process simulations more efficiently by combining dimensionality reduction with multi-fidelity Bayesian optimization. The approach uses cheap approximations alongside expensive detailed simulations, intelligently deciding when to use each, reducing the total computational cost while maintaining solution quality.

efficiencyapplications
trainingefficiencyevaluation

AdamX: Cosine similarity meets gradient descent

Sep 10, 2026

Francisco Caldas, Ruben Belo, Cláudia Soares

AdamX offers a practical alternative to standard optimizers like Adam by using cosine similarity to adaptively scale updates, potentially improving training stability and convergence speed without requiring extensive hyperparameter tuning.

AdamX is a new optimizer that combines cosine similarity with gradient descent to control how much model weights change during training. It includes a variance smoothing technique for early training stages and works across different model types and datasets with minimal setup required.

trainingefficiency

RetroThinker: Enabling Retrospective Thinking in Speech LLMs

Sep 10, 2026

Yi-Jen Shih, Puyuan Peng, Abdelrahman Mohamed et al.

Speech LLMs can now self-correct their reasoning mid-conversation without breaking real-time interaction, bridging the gap between fast speech processing and accurate complex reasoning.

RetroThinker enables speech-based AI models to improve reasoning accuracy by allowing them to revise and correct their thinking process in real-time while listening to users speak.

reasoningtrainingefficiency

Model-Aware Schedules Improve Generation via Fiberwise Optimal Transport

Sep 10, 2026

Luyi Jia, Boyan Zhang, Yilun Liu et al.

Model-aware diffusion schedules that adapt to prediction difficulty outperform fixed schedules, and the optimal allocation patterns are surprisingly consistent across different models and datasets—enabling a universal template that works without per-model tuning.

This paper improves diffusion and flow-matching models by designing schedules (signal/noise mixing ratios) that account for how well the model actually predicts at different stages.

efficiencytrainingevaluation

Near-Optimal Reinforcement Learning with Multi-Step Transition Lookahead

Sep 10, 2026

Corentin Pla, Hugo Richard, Marc Abeille et al.

Even though multi-step lookahead planning is computationally hard in the worst case, you can still plan near-optimally in polynomial time—and learning from this lookahead achieves the same regret bounds as standard RL without it.

This paper studies reinforcement learning where agents can preview the next ℓ states before choosing actions. While exact planning with this lookahead is NP-hard, the authors prove near-optimal planning is still possible in polynomial time using a randomized approximation scheme. They extend this to unknown environments and show the algorithm achieves regret matching standard RL methods.

reasoningefficiencytraining

SpecGuard: Inference-Time Backdoor Detection For Free

Sep 10, 2026

Rui Wen, Ahmed Salem, Andrew Paverd et al.

You can detect triggered backdoors in LLMs for free by monitoring an existing inference optimization (speculative decoding), without adding computation or making assumptions about trigger types.

SpecGuard detects hidden backdoors in large language models during inference by monitoring speculative decoding—a speed optimization technique. When a backdoor is triggered, the target model's behavior shifts while the draft model doesn't, causing token acceptance rates to change. This detection happens automatically without slowing down inference.

safetyefficiencyevaluation

Predicting Privacy Leakage from Weight Spectral Density

Sep 10, 2026

Richard J. Preen, Jim Smith

Weight spectral metrics provide a computationally cheap alternative to shadow model attacks for assessing privacy risk, enabling large-scale privacy audits of machine learning models.

This paper shows that spectral properties of neural network weights—like stable rank and log alpha-norm—can predict privacy leakage risk without expensive shadow model training. By analyzing weight spectra, researchers found these metrics correlate with membership inference attack success better than traditional overfitting measures, offering a faster way to audit model privacy.

safetyevaluationefficiency

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

Sep 8, 2026

Zhou Yu, Bin Bi, Shiva Kumar Pentyala et al.

When fine-tuning smaller models with expert demonstrations, don't copy entire expert trajectories—instead use on-policy correction to fix only failing steps in the model's own rollouts.

This paper shows how to improve smaller AI models on enterprise tasks by co-evolving two things: the system prompt and tool setup (harness) around the model, and the model's weights through fine-tuning. The key finding is that naive imitation learning fails because smaller models copy expert strategies they can't execute.

trainingagentsefficiency