ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2567 papers11 this month12 topics
AllTraining 46Reasoning 42Efficiency 33Evaluation 32Agents 27Architecture 18Multimodal 17Applications 16Data 9Safety 9scaling 5Alignment 2

Oct 5 – Oct 11(2)

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Oct 5, 2026

Benhao Huang, Chufan Shi, Junlin Chen et al.

Looped models can be dramatically more efficient by recognizing that recurrent states converge to fixed points, enabling truncated backpropagation, KV cache sharing, and faster RL—with learned depth priors and orthogonal injection providing better supervision than existing approaches.

This paper optimizes looped language models—models that process information through multiple recurrent passes—by leveraging fixed points in recurrent states.

trainingefficiencyarchitecture

Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model

Oct 5, 2026

Sahil Mahendrakar

Knowledge distillation can compress speech synthesis models by 10x with minimal quality loss by separating the text-to-features and features-to-audio tasks and training them independently against a frozen teacher.

Paradee is a tiny text-to-speech model created by distilling a larger 82M-parameter teacher into just 8M parameters while keeping the same voice quality. The authors use a two-stage training approach: first synthesizing training data with the teacher model, then training separate text and audio components before combining them.

Sep 28 – Oct 4(21)

Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis

Oct 2, 2026

Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan et al.

Decoder expressivity matters: simpler decoders with latent-space objectives produce better transferable geometric representations than complex pixel-space decoders, even in self-supervised settings.

This paper shows that Novel View Synthesis can learn strong 3D geometric representations if you constrain the decoder and use latent-space reconstruction instead of pixel-level targets. The authors introduce SNAP, which learns viewpoint-invariant features useful for localization, pose estimation, depth, and robot tasks—without needing explicit 3D supervision.

architecture

RNADyn: A Benchmark for Generating and Understanding RNA Dynamics

Oct 2, 2026

Yiming Huang, Lennart Bastian, Hanqun Cao et al.

A unified deep learning approach can both generate realistic RNA dynamics trajectories and predict dynamics fingerprints from static structures, bridging two previously separate tasks and improving physical accuracy through explicit physical constraints.

This paper introduces RNADynBench, a large-scale benchmark of 2,585 RNA molecular dynamics simulations, and RNADynNet, a unified model that generates realistic RNA trajectories and extracts dynamics information from single structures.

Sep 21 – Sep 27(7)

EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models

Sep 25, 2026

Kunxiong Zhu, Zhihao Shu, Hangyu Zheng et al.

Multimodal LLM serving requires rethinking GPU resource allocation around the Encode stage—treating it as a bottleneck control point rather than a separate service unlocks significant throughput gains.

EAServe optimizes serving multimodal LLMs by treating the Encode stage as a control point for the three-stage Encode-Prefill-Decode pipeline. It uses adaptive micro-batching, partial offloading, and GPU partitioning to balance resource utilization across stages, achieving 4.3x higher throughput than existing systems under latency constraints.

efficiencymultimodalarchitecture

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

Sep 24, 2026

Xinyue Zeng, Jiawei Zhang, Yujun Yan et al.

Long-horizon reasoning failures in LLMs stem from structural biases in the reasoning space itself, not just model capacity—and injecting geometric structure into the reasoning process can dramatically improve performance on hard problems.

This paper addresses why large language models struggle with long-horizon reasoning tasks by identifying two key problems: exploration bias (getting stuck in locally plausible but structurally weak paths) and compounding bias (small errors accumulating over many steps).

Sep 14 – Sep 20(12)

Benchmarking World Models for Continual Learning on Compositional Tasks

Sep 18, 2026

Haoyu Zhou, Joe Watson, Anson Lei et al.

Modular world model architectures better balance knowledge reuse with avoiding catastrophic forgetting in continual learning, but the field still lacks methods that effectively retain and reuse knowledge across sequential robot tasks.

This paper creates a benchmark to test how well world models (AI systems that learn to predict environment dynamics) can learn continuously across robot tasks without forgetting previous knowledge. The key innovation is using compositional tasks—where new tasks combine elements from earlier ones—to isolate what knowledge gets reused versus forgotten.

trainingevaluationarchitecture

Abstention and Noise Filtering: Two Missing Primitives of Softmax Attention

Sep 18, 2026

Richard Zhe Wang

Attention heads need both the ability to abstain from attending and to filter noise from values—their importance shifts with model scale, suggesting future architectures should support both primitives.

This paper identifies two missing capabilities in standard softmax attention: abstention (allowing heads to output nothing instead of always producing weighted combinations) and noise filtering (suppressing interference from mixed features).

Sep 7 – Sep 13(4)

Distance generalization in transformers: why bother with positional encoding?

Sep 10, 2026

Daniel Henrik Nevermann, Claudius Gros

Positional encodings like RoPE and ALiBi don't automatically help transformers generalize to unseen token distances—data diversity and task structure matter more than the encoding scheme itself.

This paper investigates how transformers generalize to different token distances between training and inference, using synthetic copy tasks. It compares positional encoding schemes (RoPE, ALiBi, no encoding) and finds that understanding distance generalization requires rethinking how we use positional information.

evaluationarchitecture

From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

Sep 10, 2026

Wenkang Wei, Yuan Fang, Renhe Jiang et al.

Language models have a critical handoff point where they transition from using query routing information to relying on internal knowledge—this happens at different layers across models and reveals how they internally organize and access information.

This paper investigates how large language models retrieve and use internal knowledge when answering questions by analyzing how different layers process query information versus stored knowledge.

Aug 31 – Sep 6(12)

UniMate: One Unified Model to Animate Diverse Skeletons

Sep 4, 2026

Linzhan Mou, Jiahui Lei, Zhiyang Dou et al.

For the first time, a single model can animate any skeleton topology (bipedal, quadrupedal, insects, etc.) from text alone—no per-character fine-tuning or reference motions needed at inference time.

UniMate is a foundation model that generates realistic motion for any 3D character skeleton from text descriptions, without needing to retrain for each new character type. It uses a specialized neural architecture that understands skeleton structure through graph-based attention mechanisms, and was trained on a diverse dataset of 13,000+ motion sequences across different creature types.

multimodalarchitecture

Reflection-aware Generative Novel View Synthesis

Sep 4, 2026

GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh

You can generate realistic novel views of mirror scenes by treating reflections as virtual views and using gated attention mechanisms—no retraining needed, just clever use of existing diffusion models.

This paper presents Ref-GeNVS, a method for generating novel views of scenes containing mirrors without requiring additional training. The key innovation is treating mirror reflections as complementary views by estimating the mirror plane and reflecting camera poses, then using a two-stage approach with special attention mechanisms to ensure reflections stay consistent during generation.

Aug 24 – Aug 30(2)

Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit

Aug 27, 2026

Yisen Xi

When deploying LLM agents in regulated environments, separate persona (instructions/tone) from execution (work/state) into different trust domains with a governed contract bridge—this lets you evolve agent behavior freely while maintaining execution auditability and data security.

This paper presents Persona-Execution Separation (PES), an architecture pattern for LLM agents in regulated organizations that need to evolve their instructions and tone freely while keeping their work auditable and traceable.

architecturesafetyagents

Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models

Aug 27, 2026

Frederik Berenz

Instead of pre-sizing neural network encoders at maximum capacity, you can start small and grow them incrementally as task complexity demands, achieving significant efficiency gains without sacrificing performance.

This paper introduces Successive Capacity Growth (SCG), a method that automatically expands Vision Transformer encoders in world models from minimal size upward, adding attention heads or layers only when needed to improve prediction accuracy.

Aug 17 – Aug 23(10)

Time-Aware Tranformer-Based Prediction Model for AECOPD

Aug 21, 2026

Weihao Qu, Ling Zheng, Dongyang Wang et al.

Transformer models with temporal awareness can predict COPD flare-ups from ventilator data alone, enabling faster detection in home settings without waiting for lab results.

This paper develops a transformer-based model to predict acute exacerbations of COPD using only respiratory data from home ventilators, avoiding delays from clinical lab tests. The model uses time-aware attention to track how symptoms change over time, showing better performance than traditional methods for early detection.

architecture

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

Aug 20, 2026

Adam Fisch, Shubhendu Trivedi, Fantine Huot et al.

Use value-of-information theory to decide when to invest in expensive model quality estimates before routing—this cuts estimation costs dramatically while maintaining routing accuracy.

This paper solves the problem of efficiently routing queries to the best AI model in a system with multiple specialists. The key challenge: estimating which model will perform best costs money (slow but accurate estimators vs. fast but noisy ones).

efficiencyagents

Aug 10 – Aug 16(18)

Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Aug 14, 2026

Zian Meng, Zhen Li, Chuanhao Li et al.

Separating explicit world state from appearance synthesis in video generation improves long-horizon consistency and enables direct control over predicted behavior without retraining the observation model.

Marionette is a world model for interactive games that separates world state prediction from appearance synthesis. Instead of directly generating pixels, it predicts explicit 3D skeletal poses and trajectories, uses a fixed geometric renderer to compute occlusion and geometry, then synthesizes realistic appearance on top. This makes long-horizon predictions more stable and controllable.

architecturereasoningagents

RecipeNet: A Hierarchical Transformer for Recipe Data

Aug 14, 2026

Pin-Yen Huang, Sachin Chhabra, Prasanth Sai Gouripeddi et al.

Hierarchical structure matters: representing recipes as nested sequences of structured steps, rather than flattened tables, lets models learn procedural dependencies and field interactions that improve performance on real-world synthesis and manufacturing tasks.

RecipeNet is a hierarchical Transformer model designed to learn from recipe data—ordered sequences of steps with structured fields—used in materials science, pharmaceuticals, and manufacturing. Unlike traditional tabular methods that flatten this data, RecipeNet captures both field interactions within steps and dependencies across steps, achieving better performance on recipe-based tasks.

Aug 3 – Aug 9(10)

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Aug 7, 2026

Youjun Zhao, Alex Warren, Gary K. L. Tam et al.

Mirror reflection generation requires modeling two distinct challenges—semantic consistency (what reflects) and geometric accuracy (how it's arranged)—which can be addressed by distilling relational knowledge and learning spatial transformations.

MirrorWorld tackles the problem of generating realistic mirror reflections in videos by teaching diffusion models to understand what scene content should appear in mirrors and how it should be spatially arranged. The method uses semantic guidance from visual foundation models and geometric transformation learning to ensure reflections are consistent with their surroundings.

architecturemultimodaltraining

PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents

Aug 7, 2026

Mohammad Amanlou, Parham Abed Azad, Farbod Davoodi et al.

Emotional significance and unresolved conflicts should shape what memories an AI agent retrieves, not just semantic similarity—this improves handling of complex, emotionally-laden scenarios.

PsychoAgent is a memory system for AI agents that mimics how humans remember—not just by topic relevance, but by emotional importance and unresolved conflicts. It separates factual and emotional memories, then uses an emotional filter to surface conflict-critical information when needed, showing better retrieval of conflict-relevant memories than standard similarity-based approaches.

Jul 27 – Aug 2(2)

Graph Neural Network Force Fields for Spin Dynamics in Metallic Magnets

Jul 30, 2026

Ali Rayat, Yunhao Fan, Gia-Wei Chern

GNNs can replace expensive electronic calculations for simulating spin dynamics in magnets by learning effective magnetic force fields, similar to how machine-learned potentials work for atomic systems.

Researchers developed a graph neural network framework that learns to predict magnetic forces in metallic magnets directly from electronic calculations. This approach eliminates expensive repeated electronic simulations during time evolution, enabling fast and accurate predictions of spin dynamics across different magnetic structures.

architectureefficiencyapplications

MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems

Jul 30, 2026

Mao-xun Huang, Jerry Wang, Yi-Cheng Lai et al.

Multi-agent systems can improve performance by dynamically adapting their internal communication structure at inference time, rather than relying on static pre-designed topologies.

MANTA is a framework that lets multi-agent AI systems automatically reorganize how they communicate and work together during execution. Instead of fixing agent roles and communication patterns upfront, MANTA monitors how agents collaborate and adjusts the team structure in real-time when needed—changing who talks to whom, agent responsibilities, and validation steps.

efficiencytrainingarchitecture
data
architecture
evaluation

Embedding Prediction Helps Image Generation

Oct 1, 2026

Sihan Xu, Ji Xie, Zilin Wang et al.

Dynamically predicting and updating conditioning embeddings at each generation step improves diffusion model efficiency and quality—you don't need to reuse the same embedding throughout the entire denoising process.

This paper proposes using predicted embeddings as dynamic conditioning signals in diffusion transformers instead of static embeddings. A separate transformer (NEPA) predicts image embeddings at each denoising step, allowing the conditioning to adapt to the current noise level. The approach achieves competitive image generation quality on ImageNet with significantly less training compute.

architectureefficiencytraining

SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation

Oct 1, 2026

Tianjiao Yu, Xinzhuo Li, Yifan Shen et al.

Representing 3D shapes as overlapping 2D slices rather than voxels dramatically reduces computational cost while improving topological correctness—a practical win for efficient 3D generation at scale.

SILSA is a 3D generation framework that uses compact sliding-window slice representations instead of expensive voxel tokens to generate high-resolution 3D shapes. By encoding cross-sections along three axes and adding topology supervision based on persistence diagrams, it achieves better structural quality while using 70% fewer tokens and running 58% faster than competing methods.

architectureefficiency

Hierarchical Continuous Diffusion Language Models

Oct 1, 2026

Hui Ren, Zihan Li, Chang Liu et al.

HC-DLM bridges discrete and continuous diffusion by coupling token generation with a shared latent trajectory, enabling better reasoning and constraint satisfaction than purely discrete or continuous approaches.

This paper proposes Hierarchical Continuous Diffusion Language Models (HC-DLM), which combines discrete token generation with continuous latent states in a single denoising process. Unlike existing approaches that treat these separately, HC-DLM uses the continuous latent as the only persistent state, reading tokens from it at each step.

architecturereasoningtraining

DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation

Oct 1, 2026

Zhengming Yu, Junkun Yuan, Haotian Yang et al.

By framing distribution matching as a classification problem with discriminators, DMAD eliminates the memory overhead of auxiliary models while maintaining or improving generation quality—enabling practical few-step visual generation.

DMAD improves fast image and video generation by training lightweight student models to match teacher distributions without needing an auxiliary model. It uses two discriminator heads to learn density ratios directly, making the process more efficient while achieving state-of-the-art quality in one-to-four-step generation across images and videos.

efficiencytrainingarchitecture

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Oct 1, 2026

Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi et al.

HGR enables sequence models to generate chemically valid molecules with perfect validity while capturing complex molecular topology, achieving top performance on generation and property prediction tasks without the computational cost of explicit higher-order encodings.

This paper introduces Higher-order Grammar Representation (HGR), a new way to represent molecules that captures complex structural features like ring systems by converting them into sequences of grammar rules. Unlike existing methods that struggle with computational overhead, HGR makes these structures compatible with standard sequence models while guaranteeing valid molecules.

architecturedataapplications

Generative Cinematographer: Composing Camera and Object Motion in 3D

Oct 1, 2026

Jiahan Zhang, Chaohao Yang, Namitha Guruprasad et al.

By lifting 2D video controls into explicit 3D space, you can resolve ambiguities in object motion and create videos where camera and object movements are geometrically consistent—a major improvement over 2D trajectory-based video control.

GenCine enables artists to control video generation by editing 3D camera paths and object motions in a scene scaffold, rather than using ambiguous 2D trajectories. The system projects these 3D controls into guidance maps that a pretrained video model learns to follow, producing videos with consistent camera-relative motion and improved geometric coherence.

multimodalarchitectureapplications

GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning

Oct 1, 2026

Yakun Zhu, Yi Bin, Yujuan Ding et al.

Separating geometry into common and residual components while routing visual learning through structured latents prevents representation collapse and improves 3D spatial reasoning from images.

This paper addresses 3D spatial reasoning from 2D images by introducing GeoLatent, which uses decomposed spatial representations (position, direction, geometry) with geometric supervision and routed optimization. The method prevents geometry collapse and ensures latents are actively used during learning, achieving state-of-the-art results on spatial reasoning benchmarks.

reasoningmultimodalarchitecture

Scaling Laws for Looped Mixture of Experts

Sep 30, 2026

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

Looped MoE models combine two orthogonal efficiency axes: recurrence increases computational depth while sparsity expands capacity, and their scaling laws enable principled design of efficient models that match much larger dense models at the same compute budget.

This paper develops scaling laws that jointly model looped transformers (which use recurrence for computational depth) and Mixture-of-Experts (which use sparsity for capacity).

scalingefficiencyarchitecture

Looped Diffusion Transformer

Sep 30, 2026

Yong Xien Chng, Tianyi Chen, Wenwen Tong et al.

Looping computation within denoising steps is more efficient than scaling model size—you can get better images with fewer parameters by iteratively refining representations through repeated block execution.

This paper proposes Looped Diffusion Transformer, which improves text-to-image generation by repeatedly processing the same Transformer blocks within each denoising step rather than making models larger.

efficiencyarchitecturescaling

How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?

Sep 30, 2026

Kirill Brilliantov, Alejandro Hernández-Cano, Emmanuel Abbé

For autonomous ML engineering, a minimal harness giving an LLM direct access to read, write, and bash commands performs as well as complex multi-agent systems—the model itself, not the infrastructure, drives performance.

This paper challenges the complexity of modern ML engineering agents by comparing elaborate multi-agent systems against a simple baseline where an LLM directly accesses code execution tools. The authors find that under equal time budgets, simpler agents perform as well as complex orchestrated systems, suggesting the LLM backbone matters far more than the surrounding machinery.

agentsreasoningarchitecture

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Sep 29, 2026

Jaewoo Jung, Hyeonseo Yu, Honggyu An et al.

Training MLLMs to internally reconstruct 3D scene geometry (even in compact form) improves spatial reasoning and 3D understanding, suggesting that learning to imagine scenes is more effective than explicit geometric supervision alone.

This paper teaches multimodal AI models to reason about 3D scenes by first imagining a compact 3D representation before answering questions. Instead of relying on detailed geometric details, the model learns to assemble a coarse 3D layout from multiple viewpoints—similar to how humans understand 3D space—then uses this mental model to answer spatial reasoning questions more accurately.

multimodalreasoningarchitecture

Breakdown of Local Denoising as Semantic Speciation

Sep 29, 2026

Guangkuo Liu, Mert Okyay, Yifan F. Zhang et al.

Semantic information explains why generative models transition from local to nonlocal generation at nearly the same time—this connection becomes sharper as models scale, suggesting a fundamental phase transition in how semantic structure emerges.

This paper investigates why generative models show two key moments happening nearly simultaneously: when samples commit to a semantic class (speciation) and when local context becomes insufficient for generation (nonlocality). The authors prove these windows are causally linked through semantic information distribution, showing nonlocality must occur within speciation under natural conditions.

reasoningscalingarchitecture

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

Sep 29, 2026

Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini

Factorizing attention across different data dimensions improves crop segmentation, but dataset construction choices (tile size, class definitions) have outsized impact on results and must be standardized for meaningful model comparisons.

This paper introduces PAtteRNS, a transformer-convolutional model for crop segmentation in satellite imagery that separately applies self-attention to temporal, spectral, and spatial dimensions. The authors also highlight critical dataset issues—flawed class groupings and incompatible tile-size variants—that undermine fair model comparison and suggest standardization is needed.

architecturemultimodalevaluation

Pretraining Latent Information Feedback Transformers with Teacher Supervision

Sep 29, 2026

Dor Tirosh, Ido Amos, Mor Geva

By adding recurrent feedback during pretraining via teacher-supervised state prediction, language models can improve reasoning and task performance without sacrificing training efficiency, suggesting that feed-forward architectures unnecessarily limit information flow.

This paper introduces LIFT, a transformer architecture that enables information to flow backward across layers during language model generation. Instead of the standard feed-forward design, LIFT uses teacher supervision during pretraining to train models to predict both the next token and a dense state representation derived from a teacher model.

architecturetrainingreasoning

Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Sep 29, 2026

Yu Xu, Yuxin Zhang, Xiao Yang et al.

Video models need different scaling strategies than language models—splitting experts by semantic role and allowing imbalanced routing produces better video generation than forcing uniform expert usage.

This paper proposes SplitMoE, a new way to scale video generation models using a split expert architecture that avoids forcing uniform expert usage.

scalingarchitectureefficiency

Telescopic Language Models

Sep 28, 2026

Zhilin Guo, Boqiao Zhang, Hakan Aktas et al.

You can train one model that works well at any depth by supervising random layer prefixes during training—no architectural tricks needed, just a smarter training objective that makes the model elastic across compute budgets.

This paper introduces Telescopic Language Models (TLMs), which train a single model that works effectively at every layer depth rather than requiring separate models for different compute budgets.

trainingefficiencyarchitecture

Neural Harmonic Measure Operator

Sep 28, 2026

Jinjin He, Sinan Wang, Yuchen Sun et al.

NHMO enables fast, geometry-aware PDE solving by learning a domain-specific kernel once, then reusing it for any boundary conditions or sources—eliminating expensive retraining for each new problem.

This paper presents Neural Harmonic Measure Operator (NHMO), a neural network approach that solves elliptic PDEs (like Laplace and Poisson equations) on domains with varying shapes.

reasoningefficiencyarchitecture

How to Loop MoE: Flatten the Experts, Untie the Attention

Sep 28, 2026

Shouren Wang, Chuang Ma, Mohsen Hariri et al.

To build better looped MoE models, flatten the expert hierarchy (more experts per layer, more passes) and give each pass independent attention—this lets tokens access more experts while maintaining computational efficiency.

This paper improves looped mixture-of-experts (MoE) models by flattening the expert structure and untying attention parameters. The key insight is that by doubling experts per layer and doubling passes through the network while keeping compute fixed, models can route tokens through more diverse experts, improving performance.

architectureefficiencytraining

Copy the Same, Distill the Difference: Initializing Linear Vision Transformers

Sep 28, 2026

Huaiyuan Qin, Muli Yang, Gabriel James Goenawan et al.

When adapting pre-trained vision models to use linear attention, directly copy MLP weights but distill attention behavior—this simple strategy closes the performance gap between efficient and standard transformers.

This paper shows how to initialize linear Vision Transformers (efficient attention models) using weights from standard Softmax ViTs. The key insight: copy the MLP layers directly since they learn general representations, but use distillation to transfer the attention mechanism since it's operator-specific.

efficiencyarchitecturetraining
reasoningarchitectureevaluation

Do Audio Language Models Hear and Read Distinctive Features Alike?

Sep 24, 2026

Yuanhao Chen, Peter Chin

Audio language models don't reliably represent the same phonetic features in the same direction across speech and text—suggesting these models may process the two modalities quite differently despite using a shared decoder.

This paper investigates whether audio language models represent phonetic features consistently across speech and text inputs. Using minimal pairs of phonemes that differ in single features, researchers measured whether the same distinctive features (like voicing) are encoded in the same direction in both modalities across 6 models, 7 features, and 15 languages.

multimodalevaluationarchitecture

Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

Sep 22, 2026

Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen

Diffusion LLMs can be dramatically accelerated by jointly optimizing memory I/O patterns in KV caching and using the model itself for draft-and-verify decoding, rather than treating these optimizations separately.

This paper presents Flash-dLLM, a system that speeds up diffusion language models (an alternative to traditional autoregressive LLMs) by optimizing how the model stores and reuses computation results during inference.

efficiencyarchitecture

Diffusion-Induced Spatial Attention Overlapping Community Detection

Sep 22, 2026

Kosti Koistinen, Vesa Kuikka, Joni Herttuainen et al.

By using diffusion-inspired spatial attention instead of local message passing, DISCO better captures long-range dependencies and community boundaries, making it useful for both static community detection and tracking structural anomalies in dynamic networks.

DISCO is a deep learning method for detecting overlapping communities in networks—groups where nodes belong to multiple communities simultaneously. It combines attention mechanisms with diffusion-based structural priors to overcome limitations of standard graph neural networks, and demonstrates practical value in cybersecurity by tracking how network structure changes over time.

architecturereasoning

WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory

Sep 21, 2026

Wangbo Yu, Kunhao Liu, Wenbo Hu et al.

By storing observations in a viewpoint-aware implicit 3D memory rather than explicit depth maps, video world models can generate longer, more consistent videos across different camera angles without running out of token budget.

WorldCrafter is a video world model that maintains consistent 3D-aware memory across different camera viewpoints and long time horizons. It uses an implicit memory system that compresses multi-view observations into tokens optimized for the requested viewpoint, enabling realistic minute-long video generation from a single image or text prompt while respecting what was seen before.

architecturemultimodalreasoning

Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization

Sep 21, 2026

Filipe Marinho Rocha, Inês Dutra, Vítor Santos Costa et al.

Out-of-distribution generalization requires exact representational equivalence to the generating mechanism, not statistical approximation—a criterion that constrains inference rather than training and explains why neural networks fail on novel entities while logic-based systems succeed.

This paper argues that models generalize beyond their training data only when they compute representations structurally equivalent to the underlying mechanism—not approximations.

reasoningevaluationarchitecture
architectureefficiencytraining

FAMOS: Feed-Forward 3D Articulation Modeling from Sparse Observations

Sep 17, 2026

Kevin Qu, Tao Sun, Massimiliano Viola et al.

By processing multiple sparse observations together rather than single images, and training with a motion-range supervision objective, the model better understands articulation without needing large labeled datasets—it generates synthetic training data procedurally.

FAMOS is a feed-forward model that predicts how articulated objects (like doors, drawers) move and which parts are movable from multiple sparse 3D views.

architecturemultimodal

JEPA-Anything: Learning Predictive Models across Different Worlds

Sep 17, 2026

Taoyong Cui, Zhongyao Wang, Xinyue Xu et al.

A single learning principle based on factorized predictions can work across radically different domains, suggesting world modeling doesn't need domain-specific architectures—and can even guide real scientific discovery.

JEPA-Anything is a unified framework for building predictive models across completely different domains—from videos to molecules to weather—using a technique called orthogonal predictive factorization. Instead of training separate models for each domain, it learns to decompose predictions into independent factors that work across vision, biology, physics, and clinical data.

architecturereasoningmultimodal

Agile-WAM: An Agile Tactile World Action Model for Contact-Rich Robot Control

Sep 17, 2026

Hanchu Zhou, Brendan Lynch, Raman Goyal et al.

By treating vision and tactile signals at different timescales in a lightweight architecture, you can build fast tactile-aware robot controllers (11.9ms latency) that outperform larger models on real contact-rich tasks.

This paper presents Agile-WAM, a tactile world action model that predicts future robot states and actions for contact-rich manipulation tasks.

architecturemultimodalefficiency

dQwen3.5: Hybrid-Attention Diffusion Language Models

Sep 17, 2026

Anton Xue, Litu Rout, Aditya Akella et al.

Hybrid architectures combining attention and RNNs are surprisingly efficient starting points for diffusion language models, reaching comparable performance with 2x faster training than full-attention models.

This paper shows that hybrid-attention language models (mixing attention and RNN layers) can be efficiently adapted into diffusion language models, which generate text in any order rather than left-to-right. The dQwen3.5 models reach the same training loss in half the tokens compared to full-attention baselines, while maintaining strong performance in parallel decoding.

architecturetrainingefficiency

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Sep 17, 2026

Haocheng Xi, Yiming Xie, Hexu Zhao et al.

Hybrid attention architectures that combine local Softmax with linear memory can dramatically accelerate video generation without sacrificing quality—enabling practical real-time video synthesis on modern hardware.

Video DeltaNet combines local attention with efficient linear memory to speed up video diffusion models. By mixing Softmax attention for fine details with a new Video Delta Attention mechanism for long-range context, it generates high-quality videos 14.5x faster than baseline models while maintaining visual quality.

efficiencyarchitecturemultimodal

RISC-V and machine learning: a survey

Sep 17, 2026

Shriman Keshri, Apparna Singh, Chinmaya Kumar Palo et al.

RISC-V offers a customizable, open alternative to proprietary chips for ML, but realizing its potential requires standardized extensions, better toolchain maturity, and specialized accelerator designs tailored to neural network workloads.

This survey examines how RISC-V, an open-source processor architecture, is being adapted for machine learning workloads. It analyzes existing implementations, software tools, and real-world applications, identifying strengths in energy efficiency and instruction extensions while highlighting challenges like fragmentation and standardization gaps.

architectureefficiency

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Sep 16, 2026

Sara Pieri, Evangelos Kazakos, Shizhe Chen et al.

Combining dense captioning with pixel-level grounding requires jointly optimizing text generation and mask selection—PANORAMA shows this can be done effectively by conditioning a segmenter on phrase representations and learning which masks correspond to each phrase.

This paper introduces PANORAMA, a vision-language model that generates detailed image captions while simultaneously grounding each phrase with pixel-level segmentation masks.

multimodalevaluationarchitecture

Cognitive Extensions for Dual-Process Language Agents: Memory and Self-Reflection in Interactive Environments

Sep 16, 2026

João Meneses dos Santos, Arlindo L. Oliveira

Self-reflection (validating and correcting actions at runtime) is more important than memory for making language agents reliable in interactive environments, but combining both works best.

This paper improves language agents for interactive tasks by adding two cognitive modules to SwiftSage: a memory system that stores and retrieves important experiences, and a self-reflection system that validates actions and fixes mistakes. Testing on ScienceWorld shows the full system performs best, with self-reflection being the most critical component for handling long-horizon tasks.

agentsreasoningarchitecture

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Sep 16, 2026

Zixi Chen, Akshay Vegesna, Samip Dahal et al.

Architectural design choices like recursive depth and boundary operators can fundamentally change scaling laws, enabling exponential compute efficiency improvements—challenging the idea that scaling laws are fixed.

This paper shows that architectural changes—specifically model growth through recursive depth (looping) and boundary operators—can improve how efficiently transformers use compute during training. A 7.4B looping model matches GPT-3 13B's performance with 20× less computation, with efficiency gains that grow larger at bigger scales.

architecturescalingefficiency

LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs

Sep 15, 2026

Thanapat Trachu, Samuele Cornell, William Chen et al.

Layer-specific compression in neural codecs reduces computational overhead by letting different quantization layers compress at their own optimal rates, improving efficiency without sacrificing quality.

This paper introduces LACE, a neural audio codec that compresses speech at different rates for each quantization layer rather than using a single compression step. By allowing each layer to have its own segmentation boundaries, LACE reduces sequence length and computational cost while maintaining audio quality.

efficiencyarchitecturetraining
reasoning
architecture

Studying Image Tokenizers as Visual Languages in Unified Multimodal Models

Sep 8, 2026

Siting Li, Zhengyang Wang, Simon Shaolei Du et al.

When building multimodal models, image tokenizer choice matters not just for image quality but for how well it helps the model learn text and vision together—and the best tokenizer for one task may not be best for another.

This paper studies how image tokenizers work in multimodal AI models by building a controlled training setup that tracks how well different tokenizers perform across text, image generation, and image understanding tasks.

multimodaltrainingarchitecture

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

Sep 8, 2026

Yankai Fu, Ning Chen, Junkai Zhao et al.

Adaptive tactile fusion and joint visual-tactile prediction enable VLA models to handle contact-rich manipulation where vision alone fails, achieving 71% success on dexterous tasks.

DeCAL is a vision-language-action model for robot hands that combines visual and tactile (touch) sensing to perform complex manipulation tasks. It uses specialized AI experts working together to understand scenes, imagine future states, and generate actions—all while explicitly modeling contact dynamics that cause visual occlusions during dexterous manipulation.

multimodalagentsarchitecture
multimodalarchitectureapplications

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

Sep 4, 2026

Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar et al.

Design documents with worked examples can be more maintainable than code for ML performance tools, since AI agents can reliably regenerate implementations from them, reducing the cost of keeping performance models up-to-date with new hardware and models.

SMART is a machine-learning performance modeling tool that replaces traditional code with natural-language design documents. AI agents regenerate the entire implementation from these docs on each update, eliminating tech debt while maintaining accuracy—validated against real systems like DeepSeek-V3 on TPUs.

efficiencyarchitectureapplications

Embedded Graph Flows for Categorical Graph Generation

Sep 4, 2026

Ethan Ma, Zihan Wang, Chris Siu Yeung Chow et al.

Learning continuous embeddings for graph categories instead of using fixed one-hot vectors improves molecular graph generation quality, as shown by EGF's superior performance on standard benchmarks.

This paper introduces Embedded Graph Flows (EGF), a generative model that creates realistic molecular graphs by learning continuous embeddings for node and edge types instead of using fixed one-hot vectors. The model uses a permutation-equivariant transformer to gradually transform noise into valid graph structures, achieving state-of-the-art results on molecular benchmarks like QM9 and ZINC250k.

architectureevaluation

LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams

Sep 4, 2026

Yoonju Sim, Federico Berto, Chuanbo Hua et al.

LLMs can discover domain-specific algorithm improvements by searching over heuristic families rather than predicting solutions directly—here achieving 71% win rate on quantum circuit optimization by aligning variable ordering with actual quantum cost rather than proxy metrics.

This paper uses large language models to discover better algorithms for ordering variables in quantum circuit design. The key challenge is that quantum circuits implementing Boolean functions need optimal variable orderings to minimize quantum cost, but existing heuristics optimize for the wrong metric.

reasoningarchitecture

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing

Sep 4, 2026

Pengxiang Zhao, Xing Li, Xianzhi Yu et al.

Multi-stream architectures like mHC provide more capacity than models actually use—most layers rely on just two streams, and late-layer stream mixing can be removed with minimal performance loss, suggesting opportunities for efficiency improvements.

This paper investigates how DeepSeek-V4-Flash uses its multi-stream residual pathways (mHC architecture). Researchers found that despite having four parallel streams, the model typically uses only two streams per layer, with minimal mixing between streams in later layers. Interventions show that early-layer mixing is critical for performance, while late-layer mixing is largely redundant.

architectureefficiencyevaluation

Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs

Sep 3, 2026

Yujie Zhang, Huiying Lan, Ehsan Aghapour et al.

For edge AI deployment, combining hierarchical operator parallelism with pipelining can achieve better energy efficiency and latency than using either strategy alone—Para-Pipe shows 11-23% energy improvements on real SoCs.

Para-Pipe is a framework that optimizes deep learning inference on edge devices (SoCs) by combining pipelining and parallel execution of neural network operations.

efficiencyarchitecture

The Natural Language Interaction Protocol and Standard for AI Agents

Sep 3, 2026

Luyi Xing, Rasit Onur Topaloglu, Ranjan Sinha et al.

NLIP is a practical interoperability standard for AI agents that lets you build agents independently and have them work together seamlessly, similar to how different web services communicate via HTTP.

This paper introduces NLIP, a standardized communication protocol that lets AI agents built with different frameworks and tools talk to each other. Like HTTP for web servers, NLIP provides a common language for agents to exchange messages, coordinate work, and access shared tools—enabling organizations to mix-and-match agent systems without rebuilding everything from scratch.

agentsarchitectureapplications

Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM

Sep 3, 2026

Sergii Kozyrev, Davyd Maiboroda

Recurrent attention layers like Gated DeltaNet are actually easier to quantize than softmax attention because their state-update mechanism naturally forgets noise; you can safely quantize everything to 4-bit if you use block scaling and calibrate per-module.

This paper shows that Gated DeltaNet layers in hybrid LLMs can be quantized to 4-bit precision without quality loss, contrary to conventional wisdom. The authors quantize all 496 linear layers of Qwen 3.8-27B to NVFP4 W4A4 and demonstrate it matches full precision across benchmarks while being 17.5GB smaller and 14-19% faster.

efficiencyarchitecture

Conditioning Degenerate Diffusion Models

Sep 3, 2026

Uğur Aydın, Tamer Başar

Degenerate diffusion models—where the noise schedule breaks down—can still be guided effectively using causal optimal transport, enabling diffusion-based generation in settings where standard score-based methods fail.

This paper addresses a fundamental problem in diffusion models: how to guide generation when the model has a singular diffusion coefficient (degenerate case) and the underlying data distributions are irregular or non-smooth. The authors use causal optimal transport theory to define loss functions that enable effective guidance with minimal assumptions about the data.

trainingarchitecture

Graph Machine: Towards Better Pretraining via Edges

Sep 2, 2026

Lintai Hou

You can build efficient language models by replacing dense attention with sparse routing through learned edges, maintaining full state size while accessing only a tiny fraction of tokens per layer.

Graph Machine introduces a sparse neural architecture that maintains linear-sized state while using dynamic routing through differentiable edges. By replacing 75% of dense Transformer layers with sparse GM layers in a 0.6B model, the approach achieves comparable or better performance while accessing only 2-4 tokens per attention head, reducing computational cost.

architectureefficiencytraining

Implementing neural network mixed-effects models in Template Model Builder (TMB)

Aug 31, 2026

Nan Zheng, Hoi Yiu Cheung, Vibhu Sharma et al.

TMB eliminates manual derivations for neural network mixed-effects models by automating gradient computation and random effect integration, making these powerful models easier to implement correctly and at scale.

This paper shows how to build neural network mixed-effects models (NMMs) using Template Model Builder (TMB), a tool that automatically handles the complex math needed to fit these models.

trainingarchitectureefficiency
architectureefficiencyscaling
architecture

Prompt-Conditioned Channel Attention for Hierarchical Feature Modulation toward Anatomy-Agnostic Segmentation

Aug 20, 2026

Mosharof Hossain, Md Rabiul Islam, Limon Halder et al.

Prompts can be integrated deep into image segmentation networks through channel-wise attention, rather than just at the end, making models better at finding anatomical structures across different medical imaging modalities.

This paper introduces PCCA, a mechanism that uses text prompts to guide how neural networks process medical images at multiple levels, improving segmentation accuracy across different body parts and imaging types. The method adaptively adjusts which image features matter most based on the prompt, achieving 10-23% improvements over standard approaches.

architecturemultimodalapplications

Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

Aug 20, 2026

Christos Koutsiaris

Designing model architecture around deployment constraints (CPU, 4-bit weights, fixed memory) from the start beats the conventional approach of building large then compressing—hybrid convolution-attention architectures can match attention-only models while being significantly faster at inference.

This paper presents Daedalus-150M, a small language model designed specifically for CPU inference by combining convolutions and attention strategically. Rather than shrinking a large model, the authors built from scratch with CPU constraints in mind, using full attention in only 6 of 18 blocks while the rest use short convolutions with fixed memory.

efficiencyarchitectureevaluation

The Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and Agents

Aug 20, 2026

Wei Lin, Tao Zhou, Zhaofei Xie et al.

Future software will be built from three components—storage, models, and agents—replacing traditional layered architecture. This shift means developers must rethink how they structure applications, with AI models handling logic and interfaces rather than hand-coded layers.

This paper argues that software architecture is undergoing a third major shift—from instruction-based (Software 1.0) and data-driven (Software 2.0) systems to context-and-reasoning-driven systems (Software 3.0).

architectureagentsreasoning

Ask Self, Ask Others: Relation Is All You Need

Aug 20, 2026

Yuting Ge, Pengju Yang, Mingkai Nie

Organizing token interactions into explicit Self and Exchange relations before computing attention can outperform standard attention while enabling faster implementations—suggesting relation structure matters more than raw attention scores.

This paper proposes Relation, a new token-mixing approach that organizes pairwise token interactions into explicit Self and Exchange relations before computing information flow, as an alternative to standard attention. The method shows better language modeling performance than attention across multiple model sizes and offers faster implementations.

architectureefficiency

Feature Evolution and Migration during Vision Transformer Training

Aug 20, 2026

Joonas Järve, Halil Ibrahim Aysel, Tarun Khajuria et al.

Features in Vision Transformers don't stay in one layer—they migrate between layers during training, especially early on, revealing a previously invisible dimension of how these models learn.

This paper tracks how individual features move and change across Vision Transformer layers during training using Sparse Autoencoders. By visualizing features over both network depth and training time, researchers discovered that features migrate between layers early in training, with movement favoring earlier layers, and that deeper layers stabilize faster than shallow ones.

trainingarchitecture

Lévy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention

Aug 19, 2026

Sotirios P. Chatzis, Loukas Papadoulas

Lévy Attention replaces softmax with a probabilistic formulation that automatically outputs calibrated uncertainty estimates alongside predictions—no extra parameters or passes needed, making it practical for high-stakes applications like patient risk ranking.

This paper introduces Lévy Attention, a new attention mechanism for time series that predicts both values and uncertainty in a single pass. Instead of using softmax, it formulates attention as a stochastic integral over a Poisson random measure, which naturally captures prediction confidence through two signals: disagreement (value spread) and evidence (compatibility mass).

architectureefficiency

Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets

Aug 19, 2026

Tate Berenbaum, Muthaiah Venkatachalam

Pipeline parallelism with pre-compiled shards lets you run 70B models on clusters of consumer AI PCs at interactive speeds by splitting computation across machines and optimizing GPU utilization through careful kernel fusion and speculative decoding.

This paper shows how to run large language models (like Llama 70B) across multiple Intel AI PCs by splitting the model into layers and running each layer on a different machine. Using OpenVINO optimization, speculative decoding, and request interleaving, they achieve near-optimal speed while serving multiple users simultaneously on modest hardware.

efficiencyarchitecturescaling

Proteus: Incremental Memory Activation for Long-Context Sequence Modeling

Aug 17, 2026

Reza Bayat, Ali Behrouz, Vahab Mirrokni et al.

Gradually expanding memory capacity during sequence processing—rather than using static memory—is a simple, cost-free improvement that works across different memory-based architectures and shows larger gains on longer contexts.

This paper introduces Proteus, a technique that improves how neural networks handle long text by gradually expanding memory capacity as more context arrives.

efficiencyarchitecture
architecturedata

Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training

Aug 14, 2026

Hanfeng Lu, Tianyu Feng, Suyi Li et al.

By overlapping independent computation phases and sharing GPU memory intelligently, you can train vision-language models 1.2–2.2× faster without needing more hardware or changing your RL algorithm.

Rollplex is a GPU runtime that speeds up vision-language model training by overlapping different computational phases.

efficiencytrainingarchitecture

Generating Benchmark Health Data Using a Tabular Diffusion Transformer

Aug 14, 2026

Hao Yan, Lisa Pilgram, Dan Liu et al.

You can now generate realistic synthetic tabular health data from multiple heterogeneous sources by standardizing them statistically first, then using diffusion transformers to learn and reproduce their patterns.

This paper presents a method for generating synthetic health data from multiple different database tables with varying structures. It works in two stages: first converting diverse tables into a standardized statistical format, then using a diffusion transformer to learn patterns and generate new synthetic tables.

dataarchitectureevaluation

LP-NAS: Linear Programming-based Neural Architecture Search

Aug 14, 2026

Abhishek Shukla, Ankur Sinha, Faiz Hamid

By formulating architecture search as a linear program using gradient and curvature information, LP-NAS finds better neural network designs 2-3x faster than standard differentiable NAS methods while achieving higher accuracy.

This paper proposes LP-NAS, a neural architecture search method that uses linear programming to find better network designs faster. Instead of randomly exploring architectures, LP-NAS uses mathematical optimization principles to guide the search, resulting in architectures that generalize better and are found more quickly than existing methods like DARTS.

architecturetrainingefficiency

TabSOM: A tabular-to-image encoding method based on self-organizing maps

Aug 13, 2026

David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara et al.

Self-organizing maps can encode tabular data more effectively than simpler dimensionality reduction by capturing feature relationships alongside values, improving both predictive performance and model interpretability.

TabSOM converts tabular data into images using self-organizing maps, preserving both feature values and relationships between features. Unlike existing methods that only encode individual feature values, TabSOM captures feature interactions as spatial patterns, enabling vision models to achieve better performance while remaining interpretable.

dataarchitectureevaluation

Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration

Aug 13, 2026

Sabin Roman, Ljupco Todorovski, Saso Dzeroski

SORT shifts equation discovery from brittle library selection to basis design: by learning sparse coefficients in well-chosen orthogonal bases, it provides a more stable intermediate representation that gracefully degrades under noise and sampling sparsity.

SORT is a machine learning technique that learns compact mathematical representations of dynamical systems from noisy, irregularly sampled data by fitting sparse coefficients in orthogonal basis expansions.

reasoningdataarchitecture

MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification

Aug 13, 2026

Daniel Perkins, John Squires, Janou Milligan et al.

Instead of training separate models for each domain, you can use an MLLM router to dynamically select which vision backbone handles each image, getting both better generalization and easier updates without retraining.

ARMDIL uses a multimodal language model to intelligently route images to the best-suited vision model (CNNs, self-supervised learners, or vision-language models) within an ensemble.

multimodalarchitectureevaluation

Symmetry-Breaking De Novo Crystal Generation via Markovian Jump Diffusion

Aug 13, 2026

Van Khoa Nguyen, Alexandros Kalousis

By reversing the physics concept of spontaneous symmetry breaking, this diffusion model generates complete crystal structures with proper global symmetries—a significant improvement over methods that only generate partial specifications.

This paper proposes a new method for generating crystal structures by using a diffusion-based model that starts from low-symmetry configurations and gradually breaks symmetries to create complete crystal specifications.

architectureapplications

Algebraic Decomposition Theory for Transformer Length Generalization

Aug 13, 2026

Andy Yang, Blerta Veseli, Corentin Barloy et al.

Transformers' ability to generalize to longer sequences depends on specific algebraic properties of the language being learned—properties that classical finite algebra misses but can be captured by extending decomposition theory to infinite groups.

This paper characterizes which regular languages transformers can generalize to longer sequences than they've seen during training. The authors develop new algebraic theory extending classical decomposition methods to handle transformers' unbounded counting abilities, providing a polynomial-time algorithm to predict length generalization on any regular language.

reasoningarchitectureevaluation

Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes

Aug 13, 2026

Aimilios Hadjiliasi, Louis Nisiotis

Small language models can efficiently run agent cognition (thinking and memory) on edge devices like Jetson boards, enabling virtual agents to operate independently in real-time without cloud latency.

This paper explores how small language models (SLMs) running on edge devices can power the cognitive processes of virtual agents in immersive worlds.

agentsefficiencyarchitecture

Jointly Predicting Courses and Grades Using a Transformer-Based Model

Aug 13, 2026

Paul Savala

Predicting course enrollment alongside grades significantly improves academic performance forecasting—jointly modeling what students take and how they'll perform is more accurate than predicting grades alone.

This paper presents TRACE, a transformer-based model that jointly predicts which courses students will take and their grades in those courses for upcoming semesters. Unlike traditional approaches that treat student history as a simple sequence, TRACE captures how courses taken concurrently within a semester affect performance, reducing prediction error by nearly 50% compared to grade-only models.

applicationsarchitectureevaluation

AVA-Encoder: Towards Agent-Native Video Representation Learning

Aug 12, 2026

Chuyue Li, Jinpeng Yu, Haozhe Wang et al.

Agents can now learn from and generate high-quality videos by working with structured knowledge graph representations instead of raw pixels, improving video generation quality by 20.7% over existing methods.

AVA-Encoder learns video representations as knowledge graphs that agents can reason about and edit. It converts videos into structured text and asset layers, then reconstructs videos from these representations. A natural-language feedback loop optimizes the encoding, enabling agents to work with cinematic-quality videos more effectively.

multimodalagentsarchitecture

A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery

Aug 12, 2026

Rafi Ibn Sultan, Chengyin Li, Yiannos Demetriou et al.

Combining neighborhood attention mechanisms with uncertainty-weighted loss functions enables accurate segmentation of tiny, low-contrast anatomical structures even with limited training data—a practical approach for clinical imaging tasks where manual annotation is expensive and ambiguous.

This paper presents NA-UNETR, a transformer-based deep learning model for segmenting the Left Anterior Descending artery in CT scans.

architectureefficiency

A Quantum Roadmap for Softmax Attention: Exact Born-Rule Analogs for Softmax Attention on the Probability Simplex

Aug 11, 2026

Eric A. F. Reinhardt, Adam J. Hauser

Softmax attention has an exact quantum mechanical interpretation where attention scores become measurement statistics and the softmax function emerges naturally from quantum probability rules, opening a path to quantum implementations of transformers.

This paper shows how softmax attention in transformers can be exactly implemented using quantum computing. The researchers map each component of attention—from computing scores to aggregating values—to quantum operations like measurements and rotations. They prove this works mathematically and verify the core logic in formal proof software.

architecturereasoning

sLTN: Structural Logic Tensor Networks

Aug 11, 2026

Davide Rinaldi, Luciano Serafini

If you're building neurosymbolic systems that need to reason about ordered or connected data, sLTN lets you write logical rules about structure (like "event A happens before event B") and train them end-to-end with neural networks.

sLTN extends Logic Tensor Networks to handle structured data like sequences and graphs by treating structural dimensions (time steps, positions, nodes) as first-class elements in the logical language. This lets you express temporal and relational constraints directly in logic while keeping everything differentiable for neural learning.

reasoningarchitecture

GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis

Aug 10, 2026

Alban Puech, Matteo Mazzonelli, Tamara R. Govindasamy et al.

Neural networks can solve critical power grid problems 30-85x faster than traditional solvers while staying physically consistent—opening the door to AI-powered grid analysis at scale.

GENCO is a neural network solver that handles power flow, optimal power flow, and state estimation for electrical grids in a single unified model. It matches or beats classical solvers like Newton-Raphson and IPOPT in speed and accuracy while maintaining physical consistency, and comes with an open-source framework and large datasets for reproducible research.

applicationsarchitectureevaluation

ArchAgent v2: A Case Study with the Data Prefetching Championship

Aug 10, 2026

Abraham Gonzalez, Raghav Gupta, Akanksha Jain et al.

Agentic AI can discover better microarchitecture designs than humans by breaking large search spaces into manageable pieces and embedding hardware constraints directly into the optimization loop.

ArchAgent v2 automatically designs data prefetching policies for computer processors using evolutionary search, beating hand-designed solutions in a competition. It introduces cascaded evolution (optimizing cache levels sequentially) and hardware-realizability feedback to handle the massive design space, achieving 3.8% performance improvement over baseline.

agentsarchitecture
agentsreasoningarchitecture

Timestep-Conditioned Transformers for Global Weather Forecasting

Aug 6, 2026

Sam Levang, Fran Bartolic, Ty Dickinson et al.

You can now configure forecast timesteps after training instead of before, letting one model handle both detailed short-range and stable long-range weather predictions without retraining.

GEM-3 is a weather forecasting model that solves a key problem: existing models must choose between short timesteps (detailed but error-prone) or long timesteps (stable but missing short-term details). This model lets you pick the timestep at inference time with a single trained model, and training on mixed timesteps makes forecasts more stable across longer periods.

architectureefficiencyreasoning

PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation

Aug 6, 2026

Elad Yoshai, Natan T. Shaked

Per-feature gating based on distribution distance enables better control in unpaired image translation, letting you preserve important structures while achieving realistic style changes without retraining.

PRISM is a new method for changing images from one style to another (like turning day photos into night) without paired training examples. Instead of using a single global control value, it learns a per-feature gate based on how far each image feature is from the target style, allowing precise control over what changes and what stays the same.

architectureapplicationstraining

Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor Embeddings

Aug 5, 2026

Hao Ding, Daniel Semchin, Paul M. Thompson et al.

Predicting a surface's intrinsic geometric properties (metric tensors) in continuous time, rather than directly predicting vertex positions or embeddings, produces more accurate and geometrically valid brain structure forecasts for clinical applications.

This paper presents MT-GNN, a graph neural network that predicts how brain structures will change over time by learning their intrinsic geometry (metric tensors) rather than directly predicting shape. Given a patient's past brain scans, the model forecasts the mathematical properties that define a surface's shape at future timepoints, then reconstructs the actual 3D surface.

architecture

Chained Recursive Language Models for Multi-Iteration Reasoning

Aug 5, 2026

Purbesh Mitra, Sennur Ulukus

Breaking complex reasoning into multiple fresh inference passes with shared artifacts lets LLMs catch and correct mistakes earlier, improving accuracy on multi-step tasks like counting, ordering, and multi-hop reasoning.

This paper proposes Chained Recursive Language Models (Chained RLM), a system where an LLM is called multiple times in sequence to solve complex reasoning tasks. Instead of trying to do everything in one long response, each call gets the original problem plus a summary of previous work, allowing the model to inspect and fix earlier mistakes before moving forward.

reasoningarchitecture

Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth

Aug 5, 2026

Arunava Majumder, Marius Krumm, Hendrik Poulsen Nautrup et al.

Shallow quantum generative models become strictly more powerful when augmented with classical randomness—a weak resource that enables long-range correlations impossible for unitary-only circuits at fixed depth.

This paper proves that adding shared classical randomness to shallow quantum circuits can generate probability distributions that require much deeper purely quantum circuits to replicate. The authors show this separation is achievable with simple local operations controlled by a single random bit, and demonstrate how measurement-based quantum computation naturally implements this approach.

architecturescalingreasoning

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Aug 4, 2026

Yang Yang, Qinyu Zhao, Mouxiang Chen et al.

By sharing backbone parameters across parallel branches with task-specific computation allocation, you can improve multimodal model performance without increasing model size or inference latency.

ParVL introduces a framework for scaling multimodal AI models by running multiple parallel vision and language processing branches that share the same core parameters. Instead of making models bigger or slower, it reuses existing components more efficiently and lets different tasks use different amounts of vision vs.

architectureefficiencymultimodal

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Aug 4, 2026

Junhao Chen, Mingjin Chen, Jingjia Mao et al.

Music tokenization design is more important than model size for text-to-music generation—a small model with the right token representation outperforms massive models with poor representations, challenging the field's scaling assumptions.

This paper investigates how music tokenization—the way music is converted into discrete symbols for language models—affects text-to-music generation quality. By fixing model size, data, and training approach while swapping only the tokenization scheme, researchers found that representation choice matters far more than model scale.

architecturedataevaluation

When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings

Aug 4, 2026

Christopher Schröder, Lukas Gienapp, Ferdinand Schlatt et al.

ALiBi positional encoding silently fails on long sequences due to numerical precision issues, degrading retrieval performance; using log-scaled distances instead of linear scaling provides the most reliable fix.

This paper discovers that ALiBi positional encoding—a popular method for handling long sequences in transformers—has a critical flaw: its linear bias scaling causes floating-point underflow, making many attention weights zero and blinding attention heads. The authors analyze this problem, test fixes, and show it hurts token retrieval tasks while barely affecting standard benchmarks.

architectureefficiencyevaluation
agentsreasoningarchitecture