ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2567 papers73 this month12 topics
AllTraining 46Reasoning 42Efficiency 33Evaluation 32Agents 27Architecture 18Multimodal 17Applications 16Data 9Safety 9scaling 5Alignment 2

Oct 5 – Oct 11(15)

One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline

Oct 5, 2026

Shih-Chen Tseng, Chih-Hsuan Chen, Ryan Yang et al.

Agentic systems with explicit constraint checking and visual critics can reliably preserve structural integrity in document layout tasks—achieving 68.6% fidelity versus 11-41% for prior methods—by factoring the problem into specialized stages rather than end-to-end generation.

This paper tackles the problem of automatically adapting flowchart diagrams to different aspect ratios (like fitting a pipeline figure into a paper column, slide, or social media format) while preserving all connections and content.

agentsapplicationsevaluation

Base Models Can Reason By Taking a Cue From Training Data

Oct 5, 2026

Sophie L. Wang, Amil Dravid, Rulin Shao et al.

Base models already contain reasoning capabilities encoded in their training data—you can unlock them by conditioning on the right token cues, without needing expensive RL fine-tuning.

This paper shows that base language models can achieve reasoning performance comparable to RL-trained models by using specific starting tokens (like "Okay" or "Alright") that trigger learned associations from training data. The authors demonstrate they can create new reasoning cues through data interventions and trace these effects back to specific document types in the training set.

Sep 28 – Oct 4(85)

Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

Oct 2, 2026

Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan et al.

Decoder expressivity matters: simpler decoders with latent-space objectives produce better transferable geometric representations than complex pixel-space decoders, even in self-supervised settings.

This paper shows that Novel View Synthesis can learn strong 3D geometric representations if you constrain the decoder and use latent-space reconstruction instead of pixel-level targets. The authors introduce SNAP, which learns viewpoint-invariant features useful for localization, pose estimation, depth, and robot tasks—without needing explicit 3D supervision.

architecture

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Oct 2, 2026

Ruihong Shen, Žiga Kovačič, Peter Kulits et al.

Strong vision models fail at understanding dynamic scenes through code generation—the gap between static and dynamic reconstruction is a major frontier for AI agents.

4DCodeBench is a benchmark that tests AI agents on reconstructing dynamic 3D scenes from video by writing graphics code. Agents must understand physics, deformation, and fluid dynamics to generate executable programs that recreate what they see. The benchmark reveals that current models struggle with complex motion even when they're good at static scene reconstruction.

trainingreasoningdata

BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance

Oct 5, 2026

Haojin Deng, Zhiping Lin, Yimin Yang

Monitoring centroid geometry during training can help detect and reduce spurious feature reliance, but attribute information remains partially recoverable—suggesting regularization alone isn't sufficient for complete bias removal.

BiasFlow is a monitoring toolkit that tracks how neural network backbones rely on spurious features (like gender in face recognition) through geometric analysis of feature centroids.

safetyevaluationtraining

Learning to Read the Contextual Tokens in Diffusion Transformers

Oct 5, 2026

Omer Dahary, Etai Sella, Hadar Averbuch-Elor et al.

Text tokens in image-generating transformers develop interpretable semantic representations of the emerging image that can be read with an LLM probe—and explicitly training to strengthen these representations improves generation quality.

This paper reveals what text tokens learn during image generation in multimodal diffusion transformers. Researchers built a tool to 'read' these hidden representations by connecting them to a language model, discovering they encode rich scene information early in generation. They then used these insights to improve image quality through a new training technique.

multimodaltraining

Recursive Video In-Context Learning for Agentic Robot

Oct 5, 2026

Wenrui Bao, Xinxin Liu, Bingxin Xu et al.

By organizing demonstration videos into navigable hierarchies rather than static prompts, agents can access task details only when needed, reducing context overhead while improving learning from single examples.

This paper presents Recursive Video In-Context Learning (RV-ICL), a method that helps robot agents learn from demonstration videos more effectively. Instead of feeding entire videos as prompts, RV-ICL organizes a single demo video into a hierarchical structure of sub-events (like grasps and releases) that the agent can navigate on-demand.

agentsreasoningmultimodal

Direct Intermediate Initialization for Tilted Diffusion Samplers

Oct 5, 2026

Gregory D. Bellchambers

Initializing diffusion samplers at intermediate timesteps using pulled-back clean-space posteriors and Gaussian bridges can dramatically improve sample quality, especially when the posterior has modes that are rare under the prior.

This paper improves diffusion-based posterior sampling by initializing the sampler at an intermediate step rather than starting from pure noise. The key insight is that Gaussian-tilted targets along the reverse process can be reformulated as weaker clean-space posteriors, with samples transported analytically via a Gaussian bridge.

efficiency

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Oct 5, 2026

Benhao Huang, Chufan Shi, Junlin Chen et al.

Looped models can be dramatically more efficient by recognizing that recurrent states converge to fixed points, enabling truncated backpropagation, KV cache sharing, and faster RL—with learned depth priors and orthogonal injection providing better supervision than existing approaches.

This paper optimizes looped language models—models that process information through multiple recurrent passes—by leveraging fixed points in recurrent states.

trainingefficiencyarchitecture

UniSlider: Perceptually Uniform Sliders for Continuous Image Editing

Oct 5, 2026

David Serrano-Lozano, Duygu Ceylan, Yannick Hold-Geoffroy et al.

Separating the UI slider from the underlying strength parameter and remapping it based on perceptual distance creates intuitive, predictable image editing interfaces that users prefer.

UniSlider makes image editing sliders feel natural by ensuring perceptual change increases smoothly and predictably as you move the slider. Current methods produce uneven results—some slider positions cause no visible change while others transform the image abruptly.

efficiencyapplications

MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

Oct 5, 2026

Haozhen Zhang, Haodong Yue, Quanyu Long et al.

Instead of pre-processing all memory upfront, MemPilot learns to make runtime decisions about memory curation, letting developers trade off accuracy against computational cost and speed based on their needs.

MemPilot is a framework that helps LLM agents manage memory more efficiently by deciding when to retrieve pre-stored information versus when to process raw conversation history on-demand.

agentsefficiencytraining

CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scaling

Oct 5, 2026

Yifan Zhang, Yutong Dai, Viraj Prabhu et al.

Self-verification through conformal methods lets web agents learn from their own reasoning about task progress, eliminating the need for expensive judge calls at deployment while improving training efficiency.

CLIFT trains web agents to complete browser tasks by having them verify their own actions through natural-language questions, creating a reusable signal that works both during training (with sparse rewards) and at test time (without expensive external judges). The method achieves state-of-the-art results on multiple web agent benchmarks and transfers across different models.

trainingagentsreasoning

PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

Oct 5, 2026

Yaohui Zhang, Binxu Li, Haoyi Duan et al.

Current multimodal models can approximate values from scientific figures but struggle with precision; providing source data instead of figures dramatically improves accuracy (90% to 97.4%) while reducing computational cost, suggesting a practical path for scientific data extraction.

PlotGround is a benchmark for evaluating how well AI models can extract numerical values from scientific figures.

evaluationmultimodaldata

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

Oct 5, 2026

Oliver Jaffe, Dane Sherburn

AI models are becoming exponentially better at the experimental research process itself (not just raw capability), with frontier models now reaching expert-level results using 2.3x less compute—a skill that could significantly accelerate AI R&D timelines.

TasteVal is a benchmark measuring how efficiently AI models design and conduct experiments to solve research problems. Rather than evaluating raw problem-solving ability, it measures 'experimental taste'—the skill to iteratively design good experiments and interpret results—by comparing how much compute a model needs versus human experts to reach the same performance level.

evaluationreasoningagents

Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors

Oct 5, 2026

Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi et al.

You can estimate sleep heart rate from accelerometer motion signals using deep learning, trading off some accuracy for broader coverage—useful for extracting cardiac insights from existing wearable data without optical sensors.

This paper presents SeqSmoother, a transformer-based model that estimates heart rate during sleep using only wrist accelerometer data, without requiring optical sensors.

trainingevaluationapplications

Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Oct 5, 2026

Sahil Mahendrakar

Knowledge distillation can compress speech synthesis models by 10x with minimal quality loss by separating the text-to-features and features-to-audio tasks and training them independently against a frozen teacher.

Paradee is a tiny text-to-speech model created by distilling a larger 82M-parameter teacher into just 8M parameters while keeping the same voice quality. The authors use a two-stage training approach: first synthesizing training data with the teacher model, then training separate text and audio components before combining them.

efficiencytrainingarchitecture

TAPDreamer: Transferable Adversarial Patches for World Action Models

Oct 5, 2026

Xuanyu Lu, Fengqing Jiang, Kaiyuan Zheng et al.

Small, fixed adversarial patches can severely degrade world action models across multiple robotic tasks by exploiting how visual encoders process information, highlighting that securing shared visual components is critical for robust robotic control systems.

This paper presents TAPDreamer, an attack method that uses small visual patches to fool world action models—AI systems that predict how robotic environments will change. Unlike previous attacks, TAPDreamer works without accessing the target model, instead using a public encoder to create a single patch that transfers across different tasks and robot policies.

safety
evaluationreasoningagents

What Should World Models Forget? Stratified Retention for Continual Adaptation

Oct 2, 2026

Nishit Anand, Ramani Duraiswami, Dinesh Manocha

World models need stratified forgetting strategies that preserve physical invariants while quickly adapting to environmental changes—standard continual learning metrics fail to capture this distinction and incorrectly reward frozen models.

This paper addresses a fundamental problem in continual learning for world models: knowing what to forget. Unlike traditional learning where correct labels stay correct, world models operate in changing environments where outdated knowledge must be discarded.

trainingevaluationreasoning

RNADyn: A Benchmark for Generating and Understanding RNA Dynamics

Oct 2, 2026

Yiming Huang, Lennart Bastian, Hanqun Cao et al.

A unified deep learning approach can both generate realistic RNA dynamics trajectories and predict dynamics fingerprints from static structures, bridging two previously separate tasks and improving physical accuracy through explicit physical constraints.

This paper introduces RNADynBench, a large-scale benchmark of 2,585 RNA molecular dynamics simulations, and RNADynNet, a unified model that generates realistic RNA trajectories and extracts dynamics information from single structures.

dataarchitectureevaluation

From Mixing to Tearing: Graph Decomposition in Decentralized Optimization via Message Passing

Oct 2, 2026

Kuangyu Ding, Gesualdo Scutari

Graph decomposition into tree blocks enables more efficient decentralized optimization by jointly designing subproblems and communication patterns, with convergence rates that explicitly depend on network topology and function properties.

This paper develops a new framework for distributed optimization over networks where agents minimize functions while only communicating with neighbors. Instead of traditional mixing-based approaches, the method decomposes the network graph into tree-structured blocks, with agents cooperatively solving subproblems via message passing.

trainingefficiency

LESSER: Post-Training Data Selection with Output-Layer Gradients

Oct 2, 2026

Lyuxin David Zhang, Eric Wong, Surbhi Goel et al.

You can select effective training data for LLMs using only output-layer gradients from forward passes, cutting computation costs by ~10× compared to full-gradient methods without sacrificing downstream performance.

This paper proposes LESSER, a method for selecting high-quality training data for large language models by using only output-layer gradients instead of full-parameter gradients. By leveraging cheaper forward passes rather than expensive backward passes, LESSER reduces computational cost by 9.7× for supervised fine-tuning while maintaining performance comparable to full-gradient selection methods.

trainingefficiencydata

Language Models that Play Chess and Explain Their Moves

Oct 2, 2026

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Language models can match expert-level performance in specialized domains by distilling knowledge from silent expert systems through iterative refinement, opening a path to explainable AI in games, robotics, and other domains with strong baseline models.

This paper presents Queen, a 4-billion-parameter chess model that combines a silent chess engine with a language model to play at Grandmaster level while explaining its moves.

reasoningtrainingapplications

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

Oct 2, 2026

Hui Chen, Xuan Qi, James Xu Zhao et al.

By separating strategy exploration from implementation and reusing prompt prefixes across evolution steps, you can achieve better optimization results while spending 50-100x less on LLM API calls.

FrugalEvo optimizes LLM-guided program evolution by pairing a powerful LLM that explores strategies with a cheaper LLM that implements them, while using cache-efficient prompting to reduce costs. It introduces Budget-Aware AUC to measure solution quality per dollar spent, achieving state-of-the-art results on optimization tasks at a fraction of the cost of competing methods.

efficiencyreasoningagents

Planning to Learn

Oct 2, 2026

Ian Osband

Cross-entropy beats policy gradients in classification because it's 'patient'—it optimizes for total error reduction across all future steps, not just immediate accuracy. A simple horizon-aware loss can capture this benefit while staying closer to principled gradient methods.

This paper reveals why cross-entropy outperforms exact policy gradients in classification despite having access to the true label. The key insight is that cross-entropy implicitly accounts for future learning steps, while exact policy gradients are myopic.

trainingreasoning

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Oct 2, 2026

Seo Hyun Kim, Sunwoo Hong, Younwoo Choi et al.

By identifying and selectively training on high-impact token decisions rather than full sequences, you can make diffusion language models learn more efficiently with less data.

This paper introduces Pivot-SD, a training method for masked diffusion language models that focuses on the most impactful decisions during text generation. Instead of training on entire sequences, it identifies 'pivot' tokens—commitments that significantly reduce uncertainty about remaining words—and trains only on those, using success/failure signals to guide learning.

trainingefficiencyreasoning

Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation

Oct 2, 2026

Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar et al.

Simulator-generated counterfactual rollouts can effectively bootstrap forecasting models for new policies before real deployment data exists, and these models improve further with minimal real-world calibration.

When deploying a new decision policy, prediction models face a cold-start problem because historical data reflects old policies, not the new one. This paper uses simulation to generate counterfactual training data by rolling out the new policy in a simulator, then tests whether models trained on simulated data transfer to real-world inventory control.

applicationsevaluation

Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

Oct 2, 2026

Junyoung Koh, Hao-Wen Dong

Linear STFT outperforms the more complex HCQT for vocal ensemble pitch estimation while reducing computational cost—simpler input representations can be more effective when properly designed.

This paper challenges the conventional use of harmonic constant-Q transform (HCQT) for multi-pitch estimation in vocal ensembles by showing that simpler linear STFT representations actually perform better while being much faster to compute.

evaluation

MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

Oct 2, 2026

Sean Culatana, Shang-En Huang, Kang Li

MRVQ enables one quantized index to serve all dimension-rate combinations by using truncatable residual quantization, reducing memory overhead by 17.8-22x compared to training separate indices, though with modest quality trade-offs.

This paper introduces MRVQ, a quantization method that compresses embeddings for vector search while supporting flexible trade-offs between embedding dimension and compression rate. A single index can be truncated in two ways—dropping quantization stages or embedding coordinates—to adapt to different memory and latency constraints without storing multiple separate indices.

efficiencyevaluation

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

Oct 2, 2026

Vladislav Gromadskii, David Li, Samson Gourevitch et al.

You can fine-tune fast diffusion generators for reward optimization without expensive reference rollouts by replacing intractable KL penalties with inverse-distillation regularization that provably bounds divergence.

IDRF is a method for fine-tuning masked discrete diffusion models (which generate sequences iteratively by predicting multiple tokens at once) to maximize rewards while staying close to a reference model. Instead of computing intractable likelihood penalties, it uses a clever regularization trick called inverse-distillation that upper-bounds the KL divergence.

trainingefficiencyreasoning

Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System

Oct 2, 2026

Rubén Manrique, Michelle Castellanos, Jorge Morales et al.

LLMs can sound authoritative about law they don't actually know; current models need expert oversight and source grounding for real legal work, especially outside the US where training data is sparse.

This paper evaluates how well large language models understand Colombian law by testing 15 models on 1,042 expert-validated questions covering ten legal areas. While models score well on multiple-choice questions (up to 90.5%), their free-text legal answers are rarely correct (max 45%), and they often sound confident while being wrong—a dangerous combination for non-experts relying on legal AI.

evaluationsafetyapplications

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Oct 1, 2026

Pengfei Li, Naufal Suryanto, Sicheng Zhang et al.

Current open-weight LLMs struggle with precise cybersecurity tool use (max 42% accuracy), but fine-tuning with verifiable rewards from this benchmark can make smaller models competitive with much larger ones.

KaliBench is a benchmark for evaluating how well language models can translate security analyst requests into executable commands for Kali Linux tools. It includes 8,504 query-command pairs across 1,642 tools and provides a verification system that checks both whether commands are syntactically correct and whether they actually run successfully, without needing to execute them during training.

evaluationagentssafety

Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents

Oct 1, 2026

Yen-Jen Wang, Haozhe Jiang, Shuying Deng et al.

Robots can improve their own performance through autonomous practice and skill refinement in simulation without updating model weights, then transfer successfully to real hardware—a practical path to reliable robot systems.

RPG is a framework that improves robot performance without retraining models by identifying skills from offline data, practicing in simulation with failure diagnosis, and refining symbolic skills and system prompts.

agentsreasoningtraining

Embedding Prediction Helps Image Generation

Oct 1, 2026

Sihan Xu, Ji Xie, Zilin Wang et al.

Dynamically predicting and updating conditioning embeddings at each generation step improves diffusion model efficiency and quality—you don't need to reuse the same embedding throughout the entire denoising process.

This paper proposes using predicted embeddings as dynamic conditioning signals in diffusion transformers instead of static embeddings. A separate transformer (NEPA) predicts image embeddings at each denoising step, allowing the conditioning to adapt to the current noise level. The approach achieves competitive image generation quality on ImageNet with significantly less training compute.

architectureefficiencytraining

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Oct 1, 2026

Sohyeon Kim, Yoonho Lee, Bo Liu et al.

Even advanced AI agents fail at retrieving papers that inspired real research (max 0.51 recall), revealing a critical gap in how models search scientific literature—this task requires something beyond current retrieval and reasoning approaches.

ScholarCatalyst is a benchmark dataset where 184 computer science researchers labeled which prior papers inspired their completed projects. The benchmark tests whether AI systems can retrieve these influential papers given only an initial research question and literature available at project start.

evaluationreasoningagents

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Oct 1, 2026

Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.

Representing 3D shapes as overlapping 2D slices rather than voxels dramatically reduces computational cost while improving topological correctness—a practical win for efficient 3D generation at scale.

SILSA is a 3D generation framework that uses compact sliding-window slice representations instead of expensive voxel tokens to generate high-resolution 3D shapes. By encoding cross-sections along three axes and adding topology supervision based on persistence diagrams, it achieves better structural quality while using 70% fewer tokens and running 58% faster than competing methods.

architectureefficiency

VISTA: A Visual Harness for Reasoning in an Interactive World

Oct 1, 2026

Qiushi Han, Keya Hu, Linlu Qiu et al.

Adding a simple visual memory system that preserves and lets models retrieve past observations dramatically improves multimodal models' reasoning in interactive visual environments—achieving perfect performance on challenging puzzle games.

VISTA is a visual harness that enhances multimodal models' ability to solve complex interactive visual tasks by giving them long-horizon vision and lossless visual memory. The system lets models directly perceive environments, store past observations, and actively retrieve them while reasoning.

agentsmultimodalreasoning

TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning

Oct 1, 2026

Jichao Jiang, Cristian McGee, El Houcine Bergou et al.

TACO cuts optimizer memory from 27.7 GB to 0.16 GB on 13B models by using a sparse, low-precision update strategy based on column-wise signs—making full-parameter fine-tuning practical on consumer GPUs without sacrificing model quality.

TACO is a new optimizer for fine-tuning large language models that dramatically reduces memory usage by storing only tiny gradient components per column instead of full optimizer state. It achieves 174× memory reduction compared to AdamW while maintaining accuracy, enabling fine-tuning of 30-32B models on a single GPU.

efficiencytraining

FERPO: Forward Entropy-Regularized Policy Optimization

Oct 1, 2026

Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

Using forward-KL instead of reverse-KL for policy fitting encourages broader exploration of high-value actions and avoids the computational cost of differentiating critics, leading to faster and more sample-efficient learning.

FERPO is a reinforcement learning algorithm that improves policies by using critic values directly rather than differentiating through them. It derives optimal target actions using entropy regularization and fits the actor to these targets using forward-KL divergence, which encourages exploring multiple high-value action modes while keeping importance weights stable.

trainingefficiencyreasoning

Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control

Oct 1, 2026

Akshay Balsubramani

Cost-augmented optimal transport on graphs can be solved exactly and efficiently using matrix exponentials instead of learned neural networks—no approximation or temporal discretization needed.

This paper solves the cost-augmented Schrödinger bridge problem on graphs exactly, without learning or time discretization. By reformulating state costs as a Feynman-Kac tilt of the reference process, the problem reduces to computing a standard bridge through alternating matrix exponentials. The method is exact, memory-efficient, and converges based on endpoint coupling alone.

reasoning

Hierarchical Continuous Diffusion Language Models

Oct 1, 2026

Hui Ren, Zihan Li, Chang Liu et al.

HC-DLM bridges discrete and continuous diffusion by coupling token generation with a shared latent trajectory, enabling better reasoning and constraint satisfaction than purely discrete or continuous approaches.

This paper proposes Hierarchical Continuous Diffusion Language Models (HC-DLM), which combines discrete token generation with continuous latent states in a single denoising process. Unlike existing approaches that treat these separately, HC-DLM uses the continuous latent as the only persistent state, reading tokens from it at each step.

architecturereasoningtraining

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Oct 1, 2026

Shuo Xing, Zilin Dai, Chengyuan Qian et al.

Mathematical reasoning in LLMs isn't a single skill but four distinct capabilities; focusing training on the 'Discovery' bottleneck (finding the right solution strategy) is more effective than generic math training.

This paper diagnoses why LLMs struggle with math by breaking down mathematical reasoning into four components (Discovery, Generation, Digestion, Execution) and shows that Discovery—finding the right approach—is the main bottleneck. The authors then propose a training method that uses these insights to improve math performance across different model sizes.

reasoningtrainingevaluation

Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning

Oct 1, 2026

Cristian McGee, El Houcine Bergou, Aritra Dutta

ZFO decouples direction selection from step-size selection in LLM fine-tuning, using gradient information plus two function evaluations to adaptively choose step sizes that often outperform fixed-step methods without the cost of full line searches.

This paper proposes ZFO, a lightweight optimization framework for fine-tuning large language models that intelligently selects step sizes by combining first-order gradient information with minimal zeroth-order function evaluations.

trainingefficiency

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Oct 1, 2026

Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.

Sparse autoencoder features combined with reinforcement learning enable interpretable, composable control over protein sequence generation—activating 3.75x more targeted features than previous steering approaches and improving predicted biological function.

IDiom is a specialized protein language model trained on 54 million intrinsically disordered protein regions (IDRs) that can generate functional sequences with precise control over biological features.

trainingapplicationsreasoning

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Oct 1, 2026

Zhengming Yu, Junkun Yuan, Haotian Yang et al.

By framing distribution matching as a classification problem with discriminators, DMAD eliminates the memory overhead of auxiliary models while maintaining or improving generation quality—enabling practical few-step visual generation.

DMAD improves fast image and video generation by training lightweight student models to match teacher distributions without needing an auxiliary model. It uses two discriminator heads to learn density ratios directly, making the process more efficient while achieving state-of-the-art quality in one-to-four-step generation across images and videos.

efficiencytrainingarchitecture

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Oct 1, 2026

Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi et al.

HGR enables sequence models to generate chemically valid molecules with perfect validity while capturing complex molecular topology, achieving top performance on generation and property prediction tasks without the computational cost of explicit higher-order encodings.

This paper introduces Higher-order Grammar Representation (HGR), a new way to represent molecules that captures complex structural features like ring systems by converting them into sequences of grammar rules. Unlike existing methods that struggle with computational overhead, HGR makes these structures compatible with standard sequence models while guaranteeing valid molecules.

architecturedataapplications

Decoding Looped Transformers Better for (Almost) Free

Oct 1, 2026

Weihao Liu, Huangjie Zheng, Tianrong Chen et al.

Looped Transformers naturally produce weak-to-strong prediction pairs across recurrent passes; contrasting them during decoding improves quality and enables halving compute with no training needed.

This paper introduces LoopCD, a training-free decoding method for looped Transformers that reuses intermediate predictions from earlier recurrent passes to guide token selection. By contrasting predictions from different loop depths, LoopCD improves reasoning accuracy (e.g., AIME scores from 61.88% to 73.33%) while cutting inference compute by 22-48% through fewer required loops.

efficiency

SoftServe: A Scalable Quasi-Newton Method for Deep Learning

Oct 1, 2026

Joohwan Ko, Tetiana Parshakova, Diana Cai et al.

Quasi-Newton methods—traditionally limited to convex optimization—can now scale to deep learning by using variational objectives to ensure positive curvature and GPU-friendly matrix operations, outperforming Adam on ill-conditioned problems.

SoftServe is a new quasi-Newton optimization method for training deep neural networks that handles the challenges of non-convex optimization and massive parameter counts.

trainingefficiency

Generative Cinematographer: Composing Camera and Object Motion in 3D

Oct 1, 2026

Jiahan Zhang, Chaohao Yang, Namitha Guruprasad et al.

By lifting 2D video controls into explicit 3D space, you can resolve ambiguities in object motion and create videos where camera and object movements are geometrically consistent—a major improvement over 2D trajectory-based video control.

GenCine enables artists to control video generation by editing 3D camera paths and object motions in a scene scaffold, rather than using ambiguous 2D trajectories. The system projects these 3D controls into guidance maps that a pretrained video model learns to follow, producing videos with consistent camera-relative motion and improved geometric coherence.

multimodalarchitectureapplications

From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation

Oct 1, 2026

Siqi Zhu, Suozhi Huang, Kaixuan Zhang et al.

When combining multiple RL-trained teachers into one student, the averaging method and optimizer choice matter more than raw gradient differences—response length weighting and precision loss can swing task performance by 2-5 percentage points.

This paper investigates how multiple teacher models transfer knowledge to a student model during on-policy distillation. The researchers found that loss averaging implicitly weights responses, Adam's optimizer smooths gradient differences, and low-precision arithmetic (BF16) masks small weight updates—with these factors significantly affecting which tasks the student learns best.

trainingefficiency

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

Oct 1, 2026

Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu

Self-repair in language models isn't adaptive compensation—it's pre-existing counterweights responding predictably to ablation. You can predict how a component will respond to intervention from its fixed weights alone.

When you disable a component in a language model, other parts often seem to compensate—a phenomenon called 'self-repair.' This paper shows it's not actually repair: it's pre-existing counterweights doing their normal job.

evaluation

Watch, Infer, Coordinate: Inferring Robot Partner Constraints for Zero-Shot Coordination

Oct 1, 2026

Suyu Ye, Zheyuan Zhang, Vaishnav Tadiparthi et al.

Observing joint behavior between two robots reveals hidden physical constraints better than observing a single constrained robot, enabling effective zero-shot coordination without explicit communication.

This paper tackles a practical robotics problem: how can one robot learn another robot's physical limitations (like broken joints or weak actuators) just by watching them work together, then use that knowledge to coordinate effectively on new tasks? The authors show that by analyzing how both robots move together, you can infer hidden constraints better than looking at just one robot alone.

agentsreasoning

AutoCompact: Learning When to Compact Context in Long-Horizon Coding Agents

Oct 1, 2026

Xuan Zhang, Longtao Zheng, Cunxiao Du et al.

Teaching agents to manage their own context through learned compaction decisions—rather than just handling overflow—improves performance on long-horizon coding tasks by 5-9% across different context window sizes.

AutoCompact trains coding agents to automatically decide when and how to compress their working context during long software engineering tasks. By learning when to discard stale exploration and what state to preserve, the agent improves its ability to solve repository-level coding problems while staying within context limits.

agentsreasoningtraining

DuoMind: Enabling Distributed Multi-Robot Coordination with Semantic Communication

Oct 1, 2026

Hanchu Zhou, Dechen Gao, Hang Wang et al.

Using semantic communication between robots—where they exchange meaningful descriptions rather than sensor data—enables better coordination on long-horizon tasks while keeping each robot's execution independent and reliable.

DuoMind is a framework that enables multiple robots to coordinate and work together on complex tasks by combining vision-language models for high-level reasoning with vision-language-action models for precise execution. Robots communicate through semantic messages rather than raw data, allowing them to share understanding of the task and environment while maintaining independent control.

agentsmultimodalreasoning

When Do Intrinsic Rewards Lead to Exploration?

Oct 1, 2026

Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett

Intrinsic rewards designed to encourage exploration can fail to find the most informative experiences—you need to explicitly measure whether an agent's history can substitute for real experience under different policies.

This paper examines when intrinsic reward signals (like prediction error or curiosity) actually lead to good exploration in reinforcement learning. The authors show that maximizing these rewards doesn't always produce the most informative experiences, propose a formal criterion based on counterfactual information, and demonstrate failures of existing methods with concrete examples.

trainingreasoning

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Oct 1, 2026

Lucheng Fu, Kejing Xia, Yiyang Wang et al.

LLM agents can significantly improve performance on knowledge-intensive tasks by learning persistent, source-specific models that evolve through repeated interaction—achieving up to 22.6 point gains over standard retrieval methods.

This paper introduces SourceLearn, a method for LLM agents to develop persistent, reusable understanding of external knowledge sources through repeated interaction.

trainingagentsreasoning

Faynt: Scaling and Optimizing Policies for Competitive Melee

Oct 1, 2026

Ali Janati, Nikita Kuzmin, Rohit Swamy et al.

Scaling, architecture choices, and post-training curricula matter more than model size alone—a smaller, optimized policy outperforms a larger pretrained one, suggesting careful design beats raw parameter count for complex game-playing tasks.

Faynt introduces transformer-based AI policies for Super Smash Bros. Melee that control all 26 characters with a single model. A 10M-parameter version wins 98.4% of same-character matches against existing AI opponents and beats a zero-delay competitor, using techniques like supervised pretraining on 840k human replays, curriculum learning, and distillation from a larger 75M model.

trainingreasoningagents

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

Oct 1, 2026

Juan S. Santillana

Don't trust tool-use benchmarks for small models—use verbatim-reproduction checks and token-probability probes to verify genuine capability before claiming tool use works.

Small language models can appear to use tools correctly on standard benchmarks while actually just memorizing training examples. This paper shows how keyword-matching tests miss real failures, proposes cheap diagnostic checks to catch false positives, and demonstrates a targeted fix that repairs tool-use ability in a 1.1B parameter model using minimal compute.

evaluationefficiency

Finetuning with Sampling: SFT Learns Better Than You Think

Oct 1, 2026

Aayush Karan, Sitan Chen, Yilun Du

SFT isn't inherently worse than RL for posttraining—the gap comes from data distribution mismatch. By reshaping data to be more on-policy before training, SFT can generalize better and forget less than strong RL baselines.

This paper shows that supervised finetuning (SFT) can match or beat reinforcement learning for posttraining if you transform the training data first. The authors use an MCMC sampling algorithm to gradually shift off-policy expert demonstrations toward on-policy trajectories that a reference model can actually learn from.

trainingdata

MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI

Oct 1, 2026

Negin Kafee Hernashki, Soumick Chatterjee

Anomaly detection benchmarks hide critical methodological choices; MIRTO shows that reported performance differences between methods often reflect evaluation setup rather than actual capability differences, and that registration alignment and threshold selection are major hidden sources of variance.

MIRTO is an evaluation protocol that makes explicit the hidden choices in unsupervised brain MRI anomaly detection—how maps are aligned, thresholds set, and metrics calculated.

evaluationsafety

Sample complexity bounds for categorical Markov random fields via Discrete Diffusions

Oct 1, 2026

Shivam Kumar, Nabarun Deb

Discrete diffusions can sample from structured categorical data with provable sample complexity guarantees, and weight-sharing neural networks learn scores more efficiently than fully-connected ones for this task.

This paper develops theoretical guarantees for sampling from high-dimensional categorical distributions using discrete diffusion models.

trainingreasoningscaling

Local Support Learning

Oct 1, 2026

Assaf Ben-Kish, Akarsh Kumar, James Glass et al.

You can prevent large models from forgetting old skills during new training by using a learnable gate that only activates weight updates when the input matches the current training distribution—no need to store old data.

This paper addresses catastrophic forgetting in large language models by treating it as a geometric problem in weight space.

trainingefficiencyalignment

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Oct 1, 2026

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing et al.

Current frontier AI models struggle with real enterprise data work: the best model scores 95+ on only 35% of tasks, revealing a major gap between text-to-SQL benchmarks and actual data agent capabilities needed for production systems.

Argo-Bench is an evaluation framework with 210 realistic data science tasks that test AI agents on enterprise-scale workflows.

evaluationagentsdata

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Oct 1, 2026

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc et al.

Spatially grounded self-distillation with synthetic data can teach multimodal models better visual reasoning that transfers to real-world tasks, without requiring human annotations or external teachers.

This paper improves multimodal AI models by having them learn from a smarter version of themselves that receives spatial hints about where to look in images. Using synthetic scenes with automatic object labels, the approach trains models to understand spatial relationships without human annotation, and surprisingly, these improvements transfer to real-world vision tasks.

trainingmultimodaldata

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

Oct 1, 2026

Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España et al.

When auditing AI models for medical text classification, use multiple explanation methods together and check their agreement—single methods can be misleading, especially when the model is uncertain about its prediction.

This paper evaluates how well different explanation methods agree when analyzing a medical text classifier (DeBERTa-v3). Using five different explanation techniques on medical abstracts, researchers found that explanations are most reliable when the model is confident, but diverge significantly when the model is uncertain.

evaluationsafety

Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

Oct 1, 2026

Chuqin Geng, Li Zhang, Haolin Ye et al.

Standard circuit evaluation metrics can systematically prefer incorrect mechanisms, making better discovery algorithms insufficient without fixing the evaluation objective itself.

This paper reveals a critical flaw in how mechanistic interpretability evaluates circuit discovery: the standard faithfulness metrics can prefer worse circuits that merely reproduce model behavior without actually capturing the underlying mechanisms.

evaluationsafety

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

Oct 1, 2026

Zilin Du, Bowen Yang, Boyang Albert Li

When scaling data selection with neural networks, the standard loss function causes poor generalization—a new loss function (PVM) that matches predicted values pointwise solves this and transfers better across datasets and model scales.

This paper addresses data selection for training large language models by proposing TESS, a framework that uses a neural network to score and select training examples. Unlike existing meta-learning approaches that assign per-sample weights, TESS uses a novel loss function (Pointwise Value Matching) that avoids optimization instability and improves generalization to new datasets and model sizes.

trainingdataefficiency

GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

Oct 1, 2026

Yakun Zhu, Yi Bin, Yujuan Ding et al.

Separating geometry into common and residual components while routing visual learning through structured latents prevents representation collapse and improves 3D spatial reasoning from images.

This paper addresses 3D spatial reasoning from 2D images by introducing GeoLatent, which uses decomposed spatial representations (position, direction, geometry) with geometric supervision and routed optimization. The method prevents geometry collapse and ensures latents are actively used during learning, achieving state-of-the-art results on spatial reasoning benchmarks.

reasoningmultimodalarchitecture

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

Oct 1, 2026

Kyochul Jang, Seohyeon Park, Ohchul Kwon et al.

Humanoid robots struggle with tool selection and coordinating manipulation with movement—even state-of-the-art models like GR00T show reduced accuracy on unseen tools and can execute tasks despite receiving unrelated instructions.

This paper introduces HumanoidToolBench, a benchmark for evaluating humanoid robots on tool use tasks that require selecting appropriate tools and coordinating manipulation with locomotion. The benchmark includes 3,100 demonstrations and tests seven policies, revealing significant gaps between tool selection and successful task completion.

evaluationagents

Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling

Oct 1, 2026

Daniel McBride, Pratik Khandagale, Cristina Garcia-Cardona et al.

Popular diffusion-based sampling methods have a fundamental limitation for multimodal distributions—they require exponentially long times to transition between well-separated modes, making theoretical convergence guarantees misleading about practical efficiency.

This paper challenges the claim that Wasserstein gradient flows and forward-only diffusion can efficiently sample complex multimodal distributions. Using tools from statistical physics, the authors prove these methods suffer from exponentially slow mixing times when modes are well-separated, regardless of theoretical convergence guarantees.

evaluationscaling

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

Oct 1, 2026

Yinheng Li, Justin Wagle

Modern LLMs are naturally good at making structured decisions from predefined options without fine-tuning, but targeted fine-tuning helps weaker models and specific tasks like routing—without degrading their conversational abilities.

This paper shows that large language models can already make categorical decisions (choosing from predefined options) without generating text, using their built-in token probabilities. The authors present LLM2Jev, a method to extract these decisions directly and optionally fine-tune models to improve decision-making on specific tasks, while keeping the model's text generation abilities intact.

trainingefficiencyapplications

Homomorphic Advantage Operator: Stabilizing Reinforcement Learning Under Fully Homomorphic Encryption Constraints

Oct 1, 2026

Abid Mohamed Nadhir, Ahmad Al Hanbali, Beggas Mounir

HAO enables stable reinforcement learning under FHE encryption by preventing polynomial approximation errors from accumulating—achieving 0% boundary violations versus 83.8% for unprotected baselines, making privacy-preserving RL practically viable.

This paper solves a critical problem in privacy-preserving reinforcement learning: when you encrypt data with Fully Homomorphic Encryption (FHE) for cloud computation, you must replace nonlinear operations with polynomial approximations, which causes training to diverge.

safetyefficiencytraining

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

Oct 1, 2026

Arman Behnam, Binghui Wang

Current memory systems can't measure the value of memories that are never retrieved.

Memory-augmented language models struggle to identify which memories are actually useful because some memories are never retrieved, making their value impossible to measure.

reasoningevaluationtraining

AI Emulation of Stochastic Sudden Stratospheric Warming with Interpretable Latent Structure

Oct 1, 2026

C. Daniel Boscu, Daniel Hernandez, Fabio Alvarez Ventura et al.

Deep learning models can be designed to both predict rare, extreme events accurately and produce interpretable internal representations that align with real physical phenomena, enabling better understanding of how AI makes decisions about complex systems.

Researchers built an AI model to predict rare weather events called Sudden Stratospheric Warming by learning from a simplified atmospheric model.

reasoningevaluationapplications

Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

Sep 30, 2026

Tian Xia, Minghao Liu, Yiqing Liang et al.

For imbalanced clinical tasks, optimizing prompts for ranking metrics (AUROC) instead of accuracy can dramatically improve model performance—up to 16 percentage points—because accuracy-based optimization fails when one class dominates the data.

This paper addresses class imbalance in clinical diagnosis by optimizing multimodal language models for AUROC instead of accuracy. The authors introduce Ranking-PE, a prompt optimization method that evaluates candidate prompts based on how well they rank positive cases above negative cases, rather than raw correctness.

evaluationmultimodalapplications

Semifactual Credit-Augmented Policy Optimization

Sep 30, 2026

Junshu Pan, Zhizhang Fu, Shulin Huang et al.

Token-level stability under prompt variations is a useful training signal for improving reasoning in LLMs—you can boost performance by penalizing tokens that change meaning when irrelevant prompt details change.

This paper identifies that language models trained with reinforcement learning are sensitive to irrelevant prompt changes, even when the problem stays the same. The authors propose SCAPO, an improved training method that assigns credit to individual tokens based on how stable they are under these prompt variations.

trainingreasoningalignment

Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text

Sep 30, 2026

Dulhan Jayalath, Oiwi Parker Jones

Brain-to-text decoders can accidentally learn from word timing patterns instead of brain activity; removing this shortcut with independent window processing makes the task genuinely harder but enables better learning from actual neural signals.

This paper reveals that a major brain-to-text decoding method was exploiting timing shortcuts from word duration patterns rather than learning from actual brain signals. By processing brain windows independently instead of jointly, the authors eliminate this shortcut and achieve better performance (36.6% word error rate) using simpler methods like prediction aggregation and language model priors.

evaluationreasoningapplications

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

Sep 30, 2026

Xinghao Chen, Xiangbo Gao, Jiongze Yu et al.

Video text editing requires balancing three competing goals: correct text, smooth motion, and unchanged background. This benchmark and dataset help measure those trade-offs and establish baselines for the community.

ViTeX-Bench is a benchmark for video scene text editing—replacing text on surfaces like signs and labels while keeping the rest of the video unchanged. It includes 387 real videos, evaluation metrics for text accuracy and visual quality, and a baseline editor that achieves strong results. This addresses a gap where video text editing lags behind image editing.

evaluationmultimodalapplications

Image Classifiers are Efficient Self-Supervised Video Representation Learners

Sep 30, 2026

Owais Iqbal, Sudipta Sarkar, Shyam Marjit et al.

You can efficiently learn video representations by repurposing pretrained image models with clever masking strategies, avoiding expensive 3D architectures and reconstruction overhead.

VideoMSN uses standard image Vision Transformers to learn video representations without 3D models or reconstruction. It treats videos as grids of frames and masks either spatial patches or temporal frames, then aligns the two views using a Siamese loss. This achieves top results on video benchmarks while needing 32-160x fewer training epochs than prior methods.

efficiencymultimodal

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Sep 30, 2026

Young-Jun Lee, Jinheon Baek, Soyeong Jeong et al.

By treating web search and problem-solving as co-evolving processes rather than separate steps, EvoDuet helps LLMs avoid getting stuck when they need external knowledge, improving discovery performance by 4-21% across scientific optimization tasks.

EvoDuet is a method that improves how AI models search the web while solving scientific problems. It co-evolves search queries and solutions together, letting the model decide when to search for new information versus reusing old documents. The system uses an inner loop to refine searches and an outer loop to generate solutions, achieving significant improvements on optimization tasks.

agentsreasoningapplications

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?

Sep 30, 2026

Razan El Mais, Ali Chehab, Ibrahim Issa et al.

When fine-tuning LLMs with differential privacy, untying input/output embeddings outperforms the standard weight-tied design and enables 60% memory savings—suggesting privacy-preserving training requires rethinking standard model architectures.

This paper investigates weight tying (sharing parameters between input and output embeddings) in large language models trained with differential privacy. The authors find that untying embeddings actually improves performance under DP-SGD, achieving up to 4.74% accuracy gains, while also enabling more memory-efficient privacy techniques.

safetyefficiencytraining

Turbo Harness: Instance-Adaptive Harness Optimization

Sep 30, 2026

Tunyu Zhang, Hao Wang, Kai Xu et al.

Adapting execution harnesses to individual task instances—rather than using a single global harness—consistently improves agent performance, and this adaptation can be automated by learning from previous optimization runs.

This paper introduces Turbo Harness, a system that automatically customizes AI agent execution frameworks (harnesses) for individual tasks by learning from past optimization runs. Instead of using one fixed harness for all tasks, it generates task-specific modifications that improve agent performance across diverse domains like interactive tasks, coding, and long-horizon planning.

agentstrainingefficiency

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

Sep 30, 2026

Ziyan Jiang, Jingbo Yang, Jiabao Ji et al.

Multimodal agents struggle to couple exploration and visual reasoning in 3D worlds—they can see anomalies or navigate, but struggle to do both together effectively, suggesting a fundamental gap in how these systems integrate action and perception.

WorldAuditBench is a benchmark for testing how AI agents find problems in 3D virtual worlds—like floating objects or walls you can walk through. It evaluates multimodal AI systems (vision-language models and vision-language-action models) on 213 anomaly detection tasks across 13 environments, measuring how well agents can explore systematically and visually identify issues.

evaluationmultimodalagents

Cogentic: Multi-Agent Orchestration for Automated Proof Discovery

Sep 30, 2026

Yang Cai, Vineet Gupta, Yanchen Jiang et al.

Multi-agent orchestration with verification loops can solve research-level math problems that single-shot generation cannot—by exploring multiple directions, catching errors through adversarial checking, and retaining progress across long reasoning horizons.

Cogentic is a multi-agent system that orchestrates teams of AI provers to tackle open research problems in mathematics and theoretical computer science. Instead of relying on single attempts, it uses an iterative loop where specialized agents explore different proof directions, verify results against each other, and build on confirmed findings stored in a persistent ledger.

agentsreasoningevaluation

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

Sep 30, 2026

Anson Y. Lam, Shuqing Li, Michael R. Lyu

Text-to-material generation can be more effective and controllable by outputting human-readable programs rather than raw images—users get both the final material and the explicit rules that construct it.

MatLoom generates realistic materials from text descriptions by creating compact, readable programs that define how layers combine to form appearance and physical properties.

multimodalapplications

Scaling Laws for Looped Mixture of Experts

Sep 30, 2026

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

Looped MoE models combine two orthogonal efficiency axes: recurrence increases computational depth while sparsity expands capacity, and their scaling laws enable principled design of efficient models that match much larger dense models at the same compute budget.

This paper develops scaling laws that jointly model looped transformers (which use recurrence for computational depth) and Mixture-of-Experts (which use sparsity for capacity).

scalingefficiencyarchitecture

DynaHarness: A Dynamic Physical Harness for Self-Evolving Robot Agents

Sep 30, 2026

Haoyuan Deng, Jiebin Liu, Tengxiao Zhang et al.

By separating semantic planning from physical execution and using failure evidence to guide targeted capability improvements, robots can achieve 4x better performance on long-horizon manipulation tasks compared to frozen policies.

DynaHarness is a system that improves robot manipulation by coupling semantic reasoning with physical execution monitoring. It uses a two-level architecture where a 'slow brain' plans high-level actions and a 'fast brain' grounds and monitors execution, refusing unsafe actions and requesting replans when needed. The system learns from failures to improve reusable capabilities.

agentsreasoningtraining

Looped Diffusion Transformer

Sep 30, 2026

Yong Xien Chng, Tianyi Chen, Wenwen Tong et al.

Looping computation within denoising steps is more efficient than scaling model size—you can get better images with fewer parameters by iteratively refining representations through repeated block execution.

This paper proposes Looped Diffusion Transformer, which improves text-to-image generation by repeatedly processing the same Transformer blocks within each denoising step rather than making models larger.

efficiencyarchitecturescaling

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Sep 30, 2026

Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé

For autonomous ML engineering, a minimal harness giving an LLM direct access to read, write, and bash commands performs as well as complex multi-agent systems—the model itself, not the infrastructure, drives performance.

This paper challenges the complexity of modern ML engineering agents by comparing elaborate multi-agent systems against a simple baseline where an LLM directly accesses code execution tools. The authors find that under equal time budgets, simpler agents perform as well as complex orchestrated systems, suggesting the LLM backbone matters far more than the surrounding machinery.

agentsreasoningarchitecture

Skill-Space Shooting for Autonomous Robot Policy Improvement

Sep 29, 2026

Zihang Rui, Renhao Wang, Haoxu Huang et al.

Robots can autonomously improve their policies by exploring corrections through reusable skills guided by foundation models, enabling scalable policy improvement without human demonstrations for each correction.

This paper presents skill-space shooting, a method that helps robots improve their policies by learning from their own failures without human demonstrations. The approach uses foundation models to guide exploration through reusable skills—short, familiar behaviors that can be composed to correct mistakes.

agents

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Sep 29, 2026

Jaewoo Jung, Hyeonseo Yu, Honggyu An et al.

Training MLLMs to internally reconstruct 3D scene geometry (even in compact form) improves spatial reasoning and 3D understanding, suggesting that learning to imagine scenes is more effective than explicit geometric supervision alone.

This paper teaches multimodal AI models to reason about 3D scenes by first imagining a compact 3D representation before answering questions. Instead of relying on detailed geometric details, the model learns to assemble a coarse 3D layout from multiple viewpoints—similar to how humans understand 3D space—then uses this mental model to answer spatial reasoning questions more accurately.

multimodalreasoningarchitecture

Breakdown of Local Denoising as Semantic Speciation

Sep 29, 2026

Guangkuo Liu, Mert Okyay, Yifan F. Zhang et al.

Semantic information explains why generative models transition from local to nonlocal generation at nearly the same time—this connection becomes sharper as models scale, suggesting a fundamental phase transition in how semantic structure emerges.

This paper investigates why generative models show two key moments happening nearly simultaneously: when samples commit to a semantic class (speciation) and when local context becomes insufficient for generation (nonlocality). The authors prove these windows are causally linked through semantic information distribution, showing nonlocality must occur within speciation under natural conditions.

reasoningscalingarchitecture

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Sep 29, 2026

Bingchen Yao, Haobo Xu, Haokun Lin et al.

Quantization errors in recurrent states don't matter equally: errors in long-lived memory and in dimensions that strongly influence outputs cause more accuracy loss, so allocating precision based on these factors enables aggressive compression without sacrificing performance.

STEPQuant is a quantization method that compresses the persistent memory states used in linear attention models. By analyzing how quantization errors affect model outputs differently across time and space, the method allocates precision strategically—giving more bits to memory that lasts longer and to state dimensions that matter most for predictions.

efficiency

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

Sep 29, 2026

Yi Pan, Haocheng Xi, Kan Zhu et al.

You can compress linear attention's recurrent state to 8-bit without significant quality loss by quantizing only at window boundaries and preserving outliers as special tokens—enabling faster inference on long sequences.

LeapQuant reduces the inference cost of linear attention models by quantizing their recurrent state to 8-bit precision while maintaining accuracy. It uses per-window quantization to limit error buildup and compensator tokens to handle outliers, achieving 2-3.7x speedups on real hardware.

efficiencytraining

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

Sep 29, 2026

Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini

Factorizing attention across different data dimensions improves crop segmentation, but dataset construction choices (tile size, class definitions) have outsized impact on results and must be standardized for meaningful model comparisons.

This paper introduces PAtteRNS, a transformer-convolutional model for crop segmentation in satellite imagery that separately applies self-attention to temporal, spectral, and spatial dimensions. The authors also highlight critical dataset issues—flawed class groupings and incompatible tile-size variants—that undermine fair model comparison and suggest standardization is needed.

architecturemultimodalevaluation

A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization

Sep 29, 2026

Jianru Shen

When evaluating LLM graph reconstruction, a single distance metric masks whether errors come from adding edges, removing edges, or both—the paper provides a mathematical framework to detect mixed editing and reveals different models have fundamentally different failure modes.

This paper analyzes how language models reconstruct graphs, proving mathematical bounds on the Wasserstein distance between original and reconstructed graph spectra. The bounds reveal whether a model only adds edges, only removes them, or does both—information hidden by standard distance metrics.

evaluationreasoning

EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation

Sep 29, 2026

Kuan-Po Huang, Haohe Liu, Puyuan Peng et al.

Emotion vectors in TTS models can be decomposed into a neutral-shift component and an emotion-specific component—controlling them separately via steering achieves much better emotion control than treating them as a single direction.

This paper improves emotional speech generation by decomposing emotion vectors into shared and residual components, then controlling them separately without retraining the model. The method, EmoRES, significantly outperforms prior vector steering approaches on multiple emotion metrics and human evaluation.

efficiencymultimodaltraining

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Sep 29, 2026

Hui Ren, Lei Fan, Henry Pao et al.

For long-video question answering, visually grounding object identity across time—not just retrieving relevant clips—is critical. GEB shows that organizing observations into entity biographies improves accuracy by 4+ points on day-long and week-long videos.

This paper solves a key challenge in long-video understanding: tracking the same physical object across hours or days of footage. The authors introduce Grounded Entity Biographies (GEB), which groups visual observations of the same object into retrievable "biographies" that preserve context.

multimodalreasoning

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Sep 29, 2026

Dor Tirosh, Ido Amos, Mor Geva

By adding recurrent feedback during pretraining via teacher-supervised state prediction, language models can improve reasoning and task performance without sacrificing training efficiency, suggesting that feed-forward architectures unnecessarily limit information flow.

This paper introduces LIFT, a transformer architecture that enables information to flow backward across layers during language model generation. Instead of the standard feed-forward design, LIFT uses teacher supervision during pretraining to train models to predict both the next token and a dense state representation derived from a teacher model.

architecturetrainingreasoning

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Sep 29, 2026

Paras Dahal, Anton Bakhtin, Taco Cohen et al.

Spending computation on structured decision-making about *how* to solve a problem—not just solving it directly—becomes increasingly valuable as agent tasks scale to longer horizons.

This paper introduces agentic meta-reasoning, a control system that helps AI agents manage long, complex tasks by making explicit decisions about which work to pursue, when to restart, and when to stop. A controller tracks progress compactly and decides next steps, while workers execute the actual task.

agentsreasoning

Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI

Sep 29, 2026

Cheng Qian, Kunlun Zhu, Beibin Li et al.

AI systems can improve other AI systems' performance by learning to build better execution environments—a form of test-time optimization that's reusable across tasks without modifying model weights.

This paper studies how an AI system (Builder) can learn to design better execution environments for another AI system (Target) without changing either model's weights. The Builder learns reusable principles called Meta-Skills from feedback on development tasks, then applies these to construct better environments for new tasks. Results show significant performance improvements across benchmarks.

agentsreasoningtraining