ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2567 papers11 this month12 topics
AllTraining 46Reasoning 42Efficiency 33Evaluation 32Agents 27Architecture 18Multimodal 17Applications 16Data 9Safety 9scaling 5Alignment 2

Oct 5 – Oct 11(3)

One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline

Oct 5, 2026

Shih-Chen Tseng, Chih-Hsuan Chen, Ryan Yang et al.

Agentic systems with explicit constraint checking and visual critics can reliably preserve structural integrity in document layout tasks—achieving 68.6% fidelity versus 11-41% for prior methods—by factoring the problem into specialized stages rather than end-to-end generation.

This paper tackles the problem of automatically adapting flowchart diagrams to different aspect ratios (like fitting a pipeline figure into a paper column, slide, or social media format) while preserving all connections and content.

agentsapplicationsevaluation

UniSlider: Perceptually Uniform Sliders for Continuous Image Editing

Oct 5, 2026

David Serrano-Lozano, Duygu Ceylan, Yannick Hold-Geoffroy et al.

Separating the UI slider from the underlying strength parameter and remapping it based on perceptual distance creates intuitive, predictable image editing interfaces that users prefer.

UniSlider makes image editing sliders feel natural by ensuring perceptual change increases smoothly and predictably as you move the slider. Current methods produce uneven results—some slider positions cause no visible change while others transform the image abruptly.

Sep 28 – Oct 4(16)

Language Models that Play Chess and Explain Their Moves

Oct 2, 2026

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

Language models can match expert-level performance in specialized domains by distilling knowledge from silent expert systems through iterative refinement, opening a path to explainable AI in games, robotics, and other domains with strong baseline models.

This paper presents Queen, a 4-billion-parameter chess model that combines a silent chess engine with a language model to play at Grandmaster level while explaining its moves.

reasoningtrainingapplications

Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation

Oct 2, 2026

Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar et al.

Simulator-generated counterfactual rollouts can effectively bootstrap forecasting models for new policies before real deployment data exists, and these models improve further with minimal real-world calibration.

When deploying a new decision policy, prediction models face a cold-start problem because historical data reflects old policies, not the new one. This paper uses simulation to generate counterfactual training data by rolling out the new policy in a simulator, then tests whether models trained on simulated data transfer to real-world inventory control.

Sep 21 – Sep 27(18)

Statistical attribute alignment for black-box generative AI via output post-processing

Sep 25, 2026

Kevin Jiang, Morgane Austern, Edgar Dobriban et al.

You can align generative AI outputs to target distributions by intelligently filtering multiple model queries—no model retraining needed—and this approach is provably optimal for large batches of outputs.

This paper addresses how to align AI-generated outputs with user-specified attribute distributions through post-processing, without modifying the model itself. The authors develop algorithms that select outputs from multiple queries to a generative model, ensuring attributes like gender or age match target distributions.

evaluationsafetyapplications

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

Sep 25, 2026

Md Shohel Arman, Igor Molybog

High-quality code documentation doesn't improve AI agents' ability to fix real bugs, suggesting that documentation quality and real-world problem-solving are decoupled—a finding that challenges assumptions about documentation's utility for coding agents.

This paper investigates whether better code documentation helps AI agents fix bugs in real repositories. The authors create a benchmark to measure documentation quality (based on whether code can be regenerated from descriptions) and an optimizer to improve it.

Sep 14 – Sep 20(21)

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Sep 18, 2026

Hongyang Du, Lan Yan, Christian Flores et al.

Procedural memory—a continuously updated library of natural-language design skills—enables frozen frontier models to improve at complex agentic tasks by learning from execution failures without model retraining or human annotation.

This paper shows how a frozen AI model can continuously improve at graphic design by building and refining a library of reusable design procedures from real user projects. Without updating the model's weights or using human labels, the system learns 139 design skills from 1,406 real briefs, improving success rates from 73% to 99% by accumulating new procedures and fixing failed ones.

agentstrainingapplications

Cross-sector generalization of accident-process role classification in occupational accident narratives

Sep 18, 2026

Aho Yapi, Pierre Latouche, Arnaud Guillin et al.

Fine-tuned language models can generalize accident report classification across different industries without retraining, enabling scalable occupational safety analysis across sectors.

This paper develops an automated system to classify key information in occupational accident reports (work situations, unsafe conditions, events, consequences) and tests whether models trained on construction-sector narratives can work across different industries like metallurgy and chemistry.

Sep 7 – Sep 13(9)

Can Edge-Deployable Vision-Language Models Identify Species?

Sep 10, 2026

William Zhou, Mayukha Siripuram, Xiao Yan et al.

Small deployable VLMs can identify species above chance but degrade sharply on real camera-trap images; specialized training data outweighs model scale, and all models occasionally hallucinate fake species names when uncertain.

This paper evaluates whether small vision-language models (2-8B parameters) suitable for edge deployment can accurately identify animal species from camera-trap photos.

evaluationmultimodalapplications

Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact

Sep 10, 2026

Masahiro Kato, Daiki Honma, Taka Kato

Businesses can now quantify the ROI of appearing in generative AI outputs by combining visibility data with causal inference, enabling marketing decisions in an AI-driven world where traditional metrics don't apply.

This paper introduces a framework to measure how generative AI systems (like ChatGPT) affect business outcomes when companies optimize their visibility in AI-generated answers.

applications

Aug 31 – Sep 6(30)

Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction

Sep 4, 2026

Sihwa Park

Embodied, tangible interaction can make complex AI processes intuitive and engaging without technical jargon—by letting people physically manipulate outputs, they experientially understand how diffusion models gradually refine noise into coherent images.

Diffusion TV is an interactive art installation that lets people physically experience how diffusion models work by manipulating a modified CRT TV's antenna to control image clarity. As visitors tune the antenna and knob, they directly engage with the denoising process—the core mechanism behind AI image generation—while exploring AI-generated animals across past, present, and future timelines.

applicationsevaluation

RegionFed: Federated Learning for Personalized Query Understanding in Heterogeneous Retail Environments

Sep 4, 2026

Quoc H. Nguyen, Ali Lafzi, Abhijeet Phatak et al.

Gradient-level personalization in federated learning works reliably across transformer architectures where parameter-level methods fail, enabling privacy-preserving retail systems that adapt to regional differences while maintaining strong performance.

RegionFed solves a critical problem in retail search: different regions have different query patterns and product preferences, but standard federated learning produces one-size-fits-all models that perform poorly everywhere.

Aug 24 – Aug 30(3)

SWE-Prime: Fewer Trajectories, Better Performance

Aug 27, 2026

Dewu Zheng, Ruizhe Ye, Yanlin Wang et al.

Filtering training data for code-fixing tasks at both trajectory and step levels beats training on all successful examples—quality and relevance of supervision matter more than quantity.

SWE-Prime improves how AI models learn to fix software bugs by being smarter about which training examples to use. Instead of training on all successful bug fixes, it filters trajectories (problem-solving paths) at two levels—first selecting high-quality complete solutions, then identifying which individual steps within those solutions are actually worth learning from.

trainingdataapplications

From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

Aug 27, 2026

Dewu Zheng, Yanlin Wang, Xiwen Wang et al.

Current LLMs struggle with multi-round code review—their performance drops significantly as review iterations increase, they miss complex defects, and they fail to track how issues change across multiple rounds of feedback.

MCR-Bench is a new benchmark for evaluating AI models on realistic code review tasks that involve multiple rounds of back-and-forth interaction between developers and reviewers.

efficiencyapplications

Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors

Oct 5, 2026

Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi et al.

You can estimate sleep heart rate from accelerometer motion signals using deep learning, trading off some accuracy for broader coverage—useful for extracting cardiac insights from existing wearable data without optical sensors.

This paper presents SeqSmoother, a transformer-based model that estimates heart rate during sleep using only wrist accelerometer data, without requiring optical sensors.

trainingevaluationapplications
applicationsevaluation

Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System

Oct 2, 2026

Rubén Manrique, Michelle Castellanos, Jorge Morales et al.

LLMs can sound authoritative about law they don't actually know; current models need expert oversight and source grounding for real legal work, especially outside the US where training data is sparse.

This paper evaluates how well large language models understand Colombian law by testing 15 models on 1,042 expert-validated questions covering ten legal areas. While models score well on multiple-choice questions (up to 90.5%), their free-text legal answers are rarely correct (max 45%), and they often sound confident while being wrong—a dangerous combination for non-experts relying on legal AI.

evaluationsafetyapplications

Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features

Oct 1, 2026

Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.

Sparse autoencoder features combined with reinforcement learning enable interpretable, composable control over protein sequence generation—activating 3.75x more targeted features than previous steering approaches and improving predicted biological function.

IDiom is a specialized protein language model trained on 54 million intrinsically disordered protein regions (IDRs) that can generate functional sequences with precise control over biological features.

trainingapplicationsreasoning

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Oct 1, 2026

Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi et al.

HGR enables sequence models to generate chemically valid molecules with perfect validity while capturing complex molecular topology, achieving top performance on generation and property prediction tasks without the computational cost of explicit higher-order encodings.

This paper introduces Higher-order Grammar Representation (HGR), a new way to represent molecules that captures complex structural features like ring systems by converting them into sequences of grammar rules. Unlike existing methods that struggle with computational overhead, HGR makes these structures compatible with standard sequence models while guaranteeing valid molecules.

architecturedataapplications

Generative Cinematographer: Composing Camera and Object Motion in 3D

Oct 1, 2026

Jiahan Zhang, Chaohao Yang, Namitha Guruprasad et al.

By lifting 2D video controls into explicit 3D space, you can resolve ambiguities in object motion and create videos where camera and object movements are geometrically consistent—a major improvement over 2D trajectory-based video control.

GenCine enables artists to control video generation by editing 3D camera paths and object motions in a scene scaffold, rather than using ambiguous 2D trajectories. The system projects these 3D controls into guidance maps that a pretrained video model learns to follow, producing videos with consistent camera-relative motion and improved geometric coherence.

multimodalarchitectureapplications

LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them

Oct 1, 2026

Yinheng Li, Justin Wagle

Modern LLMs are naturally good at making structured decisions from predefined options without fine-tuning, but targeted fine-tuning helps weaker models and specific tasks like routing—without degrading their conversational abilities.

This paper shows that large language models can already make categorical decisions (choosing from predefined options) without generating text, using their built-in token probabilities. The authors present LLM2Jev, a method to extract these decisions directly and optionally fine-tune models to improve decision-making on specific tasks, while keeping the model's text generation abilities intact.

trainingefficiencyapplications

AI Emulation of Stochastic Sudden Stratospheric Warming with Interpretable Latent Structure

Oct 1, 2026

C. Daniel Boscu, Daniel Hernandez, Fabio Alvarez Ventura et al.

Deep learning models can be designed to both predict rare, extreme events accurately and produce interpretable internal representations that align with real physical phenomena, enabling better understanding of how AI makes decisions about complex systems.

Researchers built an AI model to predict rare weather events called Sudden Stratospheric Warming by learning from a simplified atmospheric model.

reasoningevaluationapplications

Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

Sep 30, 2026

Tian Xia, Minghao Liu, Yiqing Liang et al.

For imbalanced clinical tasks, optimizing prompts for ranking metrics (AUROC) instead of accuracy can dramatically improve model performance—up to 16 percentage points—because accuracy-based optimization fails when one class dominates the data.

This paper addresses class imbalance in clinical diagnosis by optimizing multimodal language models for AUROC instead of accuracy. The authors introduce Ranking-PE, a prompt optimization method that evaluates candidate prompts based on how well they rank positive cases above negative cases, rather than raw correctness.

evaluationmultimodalapplications

Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text

Sep 30, 2026

Dulhan Jayalath, Oiwi Parker Jones

Brain-to-text decoders can accidentally learn from word timing patterns instead of brain activity; removing this shortcut with independent window processing makes the task genuinely harder but enables better learning from actual neural signals.

This paper reveals that a major brain-to-text decoding method was exploiting timing shortcuts from word duration patterns rather than learning from actual brain signals. By processing brain windows independently instead of jointly, the authors eliminate this shortcut and achieve better performance (36.6% word error rate) using simpler methods like prediction aggregation and language model priors.

evaluationreasoningapplications

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

Sep 30, 2026

Xinghao Chen, Xiangbo Gao, Jiongze Yu et al.

Video text editing requires balancing three competing goals: correct text, smooth motion, and unchanged background. This benchmark and dataset help measure those trade-offs and establish baselines for the community.

ViTeX-Bench is a benchmark for video scene text editing—replacing text on surfaces like signs and labels while keeping the rest of the video unchanged. It includes 387 real videos, evaluation metrics for text accuracy and visual quality, and a baseline editor that achieves strong results. This addresses a gap where video text editing lags behind image editing.

evaluationmultimodalapplications

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Sep 30, 2026

Young-Jun Lee, Jinheon Baek, Soyeong Jeong et al.

By treating web search and problem-solving as co-evolving processes rather than separate steps, EvoDuet helps LLMs avoid getting stuck when they need external knowledge, improving discovery performance by 4-21% across scientific optimization tasks.

EvoDuet is a method that improves how AI models search the web while solving scientific problems. It co-evolves search queries and solutions together, letting the model decide when to search for new information versus reusing old documents. The system uses an inner loop to refine searches and an outer loop to generate solutions, achieving significant improvements on optimization tasks.

agentsreasoningapplications

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

Sep 30, 2026

Anson Y. Lam, Shuqing Li, Michael R. Lyu

Text-to-material generation can be more effective and controllable by outputting human-readable programs rather than raw images—users get both the final material and the explicit rules that construct it.

MatLoom generates realistic materials from text descriptions by creating compact, readable programs that define how layers combine to form appearance and physical properties.

multimodalapplications

Retrieving Biblical Intertextual References in Karen Blixen's Seven Gothic Tales

Sep 28, 2026

András Kovács, Alexander Conroy, Daniel Hershcovich et al.

Fine-tuned dense retrievers can effectively find biblical allusions and paraphrases in literature, but computational models work best as collaborative tools with human experts rather than as autonomous discovery systems—some high-ranked predictions scholars deem meaningful fall outside official...

This paper tackles computational retrieval of biblical references in literary texts, using Karen Blixen's Seven Gothic Tales as a case study.

applications

Scaling Long-Form Story Generation via Narrative State Tracking

Sep 28, 2026

Zhennan Wan, Jianfei Chen

Explicitly tracking narrative state (characters, events, plot requirements) as a structured agent task enables LLMs to write coherent long-form stories without special training, scaling from short stories to novel-length works.

This paper introduces NstAgent, a framework that helps large language models write longer stories by tracking narrative elements like characters and plot points. The system maintains consistency across 10K-100K word stories without degrading quality, addressing a key challenge in scaling creative writing to full-length novels.

agentsreasoningapplications

FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

Sep 28, 2026

Hoyoung Lee, Suyeol Yun, Jack Haverty et al.

Automating rubric generation with expert oversight lets you evaluate complex financial AI systems at scale without manually writing new rubrics for each task, while maintaining quality comparable to human expert grading.

FinAutoRubric automates the creation of evaluation rubrics for financial research agents by combining expert guidance with AI-generated, task-specific criteria.

evaluationapplicationsagents
evaluationagentsapplications

Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach

Sep 25, 2026

Damiano Brigo, Raphaël Huser, Dan Leonte

Neural calibration of financial models should quantify posterior parameter uncertainty and propagate it through pricing models—point estimates alone can be materially unreliable for exotic derivatives, and information-theoretic explainability reveals which market data regions constrain which pa...

This paper develops a neural framework for calibrating rough Heston volatility models that captures parameter uncertainty rather than just point estimates. It combines simulation-based inference with neural surrogates to produce uncertainty-aware price intervals for exotic options, and introduces an explainability method to identify which market observations drive parameter learning.

evaluationapplications

DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

Sep 25, 2026

Quang Nguyen, Hieu Nguyen, Hien Hoang et al.

Building effective AI tutors for non-English regions requires both technical efficiency (faster inference on consumer hardware) and domain-specific learning (accumulating local knowledge from real interactions rather than relying on pre-training).

DeepEdu-v1 is an AI tutoring system for Vietnamese students that runs locally to protect data privacy and avoid hallucinations from Western-trained models. It uses two key innovations: a smarter way to handle long conversations that reduces processing time by 35%, and a self-improving system that learns from past tutoring interactions instead of requiring expensive retraining.

efficiencyapplicationsagents

RAPID: Robot Agentic Programming from Demonstrations

Sep 24, 2026

Yuyao Liu, Jiayuan Mao, David Hsu et al.

By combining visual demonstrations with agentic code refinement and object-centric representations, RAPID enables robots to learn generalizable manipulation skills that transfer across object variations and scene configurations.

RAPID automatically generates reusable robot programs from a single human video demonstration. It uses AI coding agents to create, test, and refine programs that work across different objects and environments by learning the underlying strategy rather than memorizing specific motions.

agentsreasoningapplications

Coding Agents for Generalized Task and Motion Planning Problems

Sep 24, 2026

Matteo Merler, Bowen Li, Josh Roy et al.

Coding agents can automatically synthesize generalizable planning programs better than traditional planners or one-shot LLM generation, suggesting LLMs writing code is a practical approach to automating robotics planning.

This paper shows that AI coding agents (like Claude and GPT models) can automatically write programs that solve robot task-and-motion planning problems across different scenarios. Given access to a simulator, agents develop reusable code within a budget, then freeze it for evaluation on new instances.

agentsreasoningapplications

A Living Benchmark for Information Retrieval from Electronic Health Records

Sep 24, 2026

Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani et al.

Automated benchmark generation validated by domain experts enables continuous evaluation of clinical AI systems, exposing real gaps (like multi-document synthesis) that static benchmarks miss.

Researchers created BRIE, an automatically-generated benchmark for testing how well AI systems retrieve and synthesize information from patient medical records.

evaluationapplicationsdata

Jev-Mobile: Jev as an Executor for Mobile GUI Agents

Sep 24, 2026

Linghua Zhang

Decoupling VLM planning from action execution using a lightweight executor reduces model serving costs dramatically without sacrificing task performance on mobile GUI automation.

Jev-Mobile improves mobile GUI agents by separating planning from execution: a vision-language model makes high-level decisions infrequently, while a lightweight decision model (Jev) handles repeated low-level actions. This cuts inference costs by 73% and execution time by 33% while maintaining 79% task success on Android tasks.

agentsefficiencyapplications

ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints

Sep 24, 2026

Sriram Kannan, Swetha Saseendran, Vishnu Vardhan Reddy Kandi et al.

Structured event graphs from legal documents improve reasoning tasks like classification and QA, but only when you already have the right documents—they don't help with initial retrieval.

ARGUS builds structured Event Knowledge Graphs from employment-discrimination legal complaints using a 5W1H schema and LLMs. The system extracts facts, organizes events with temporal and causal relationships, and merges them into document-level graphs. Testing shows these graphs improve claim classification and help answer legal questions when relevant documents are already retrieved.

dataapplicationsreasoning

Search-Aware Reinforcement Learning for Multi-Component Query Understanding in Roblox Game Search

Sep 24, 2026

Nayoung Choi, Shengjian Chen, Xiaokai Wei et al.

Optimizing search query understanding components individually with search-engine-derived rewards outperforms single end-to-end optimization, showing that understanding how each component affects downstream retrieval matters more than just matching labels.

This paper presents a reinforcement learning framework for query understanding in search systems that optimizes multiple components (like intent classification and query expansion) separately using rewards from live search engine interactions, rather than treating it as a single end-to-end problem.

trainingreasoningapplications

GridSFM: A Foundation Model for Solving AC Optimal Power Flow

Sep 24, 2026

Luke Bhan, Weiwei Yang, Margaret Capetz et al.

Foundation models can solve complex physics optimization problems across different system sizes by combining pretraining on diverse topologies with physics-informed fine-tuning, enabling practical deployment on real power grids without retraining from scratch.

GridSFM is a foundation model that solves AC Optimal Power Flow (a critical power grid optimization problem) by pretraining a 15M-parameter graph neural network across 54 different grid topologies, then fine-tuning it with physics-informed methods.

reasoningapplications

FleXray: Universal Clinical X-ray Segmentation

Sep 22, 2026

Victor Ion Butoi, Vivek Gopalakrishnan, John V. Guttag et al.

You can train powerful medical imaging models without expensive manual annotation by simulating realistic training data from existing 3D datasets—FleXray segments full-body X-rays and works on real clinical data despite being trained entirely on synthetic images.

FleXray is a generalist AI model that segments 60 anatomical structures in clinical X-rays across the entire body. Rather than manually labeling thousands of X-rays, the researchers built a physics-based simulator that generates realistic synthetic X-rays from existing 3D CT scans, then trained the model on these simulations.

trainingdataapplications

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

Sep 22, 2026

Om Nepal, Sushant Aryal, Oluseyi Olukola et al.

Don't use compile rate to evaluate LLM vulnerability repair; it's gamed by non-repairs and dominated by evaluation setup, not model quality. Use change-aware metrics like diff_F1 as a first-pass screen before running actual tests.

This paper shows that compile rate—a common metric for measuring LLM-based code vulnerability repair—is unreliable because it's dominated by dataset artifacts rather than model quality, shifts dramatically with compiler flags, and even rewards broken patches.

evaluationsafetyapplications

Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows

Sep 22, 2026

Remy Stewart, Olabode Anise, Andrew Hogan et al.

AI design tools deliver real time savings (~20%) in controlled settings, but benefits vary by user expertise—product managers gain more than professional designers, indicating task-dependent value.

Researchers tested whether AI-powered design tools (Figma Make) actually save time by having 50 designers and 50 product managers complete design tasks with and without the tool. They found about 20% faster completion times with AI assistance, especially for non-designers, suggesting these tools may let product managers do more design work themselves.

evaluationapplicationsefficiency

LoRA-generating hypernetworks for efficient on-device LLM generative personalization

Sep 21, 2026

Sean Augenstein, Li Ding, Jihwan Lee et al.

Hypernetworks can efficiently generate personalized LoRA adapters on mobile devices by mapping user context to weight modifications, avoiding the latency costs of context extension while remaining computationally feasible for resource-constrained devices.

This paper presents a method for personalizing language models on mobile devices by training a hypernetwork that generates customized LoRA (low-rank adaptation) weights based on a user's context.

efficiencyapplications

Harness-Zero: Harness Distillation via Agent-as-Harness

Sep 21, 2026

Haoran Ye, Yuxing Lu, Haonan Dong et al.

You can distill specialized harness behaviors into model weights by having an intermediate agent translate between different harness action spaces during training, letting you deploy with simpler harnesses while keeping performance gains.

This paper tackles how to transfer the benefits of specialized agent harnesses (external systems that improve model-environment interaction) into model weights so they work with simpler harnesses at deployment.

trainingagentsapplications

Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks

Sep 21, 2026

Muzhe Wu, Zuchen Li, Xu Wang et al.

Contextualizing visual instructions to match a user's real workspace—through AI-generated images and videos—improves task performance and trust compared to generic pre-authored tutorials.

This paper presents a system that generates live visual instructions tailored to a user's specific workspace and task progress, rather than showing pre-recorded tutorials. Using AR and generative AI, it creates goal images and demo videos that match the user's actual environment, helping them complete physical tasks more accurately and confidently.

applicationsmultimodalevaluation

JAREX: An Acquisition Function for Multi-Objective Algorithmic Process Characterization

Sep 21, 2026

Xinyang Li, Kevin Stone, Ajit Vikram

For multi-objective process characterization, JAREX reduces experimental costs by 50%+ compared to traditional design-of-experiments approaches by adaptively focusing on the boundaries where products transition from acceptable to unacceptable.

JAREX is a machine learning method that helps pharmaceutical companies efficiently map out which combinations of process parameters produce acceptable products. Instead of running many traditional experiments, it uses Bayesian optimization to intelligently select which experiments to run next, cutting the number of tests needed in half while maintaining accuracy.

applications
trainingevaluationapplications

Available Guardrails: Certifying Selective Prediction across ML Systems

Sep 18, 2026

Parivesh Priye, Yufeng Wang, Haibin Ling et al.

When deploying safety-critical ML systems with selective prediction, finite calibration data is the bottleneck—smart partition selection and error budget reallocation can recover 60% of the theoretical coverage gain, but naive approaches recover almost none.

This paper addresses how to safely deploy selective predictors (models that abstain when uncertain) by certifying they meet precision targets for specific user groups. The key challenge is that with limited calibration data, some groups may not have enough evidence to certify safety.

safetyevaluationapplications

Gricea: An Open Science Platform for Conversational AI Research

Sep 18, 2026

Nikhil Sharma, Yunlin Gong, Xinyang Cheng et al.

Standardized, shareable research artifacts dramatically improve reproducibility in conversational AI research—Gricea enables researchers to replicate 93% of studies and identify gaps that would otherwise block replication.

Gricea is an open-science platform that lets researchers package conversational AI studies as reusable, shareable artifacts. It standardizes how study procedures, chatbot systems, and conversation tasks are documented and deployed, making it easier to replicate and build on prior research.

evaluationapplications

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Sep 17, 2026

Ji Xie, Dewei Zhou, Xinyu Huang et al.

By combining real image supervision with pure-color anchors and a shared hex-prompt interface, you can now control exact object colors in both image generation and editing tasks with professional-grade precision.

Paint-Anything enables precise color control in AI image generation and editing by letting users specify exact colors using hex codes (like #FF5733). The system learns to understand hex values through a training dataset of 500K images with color labels and pure-color reference images, achieving much better color accuracy than existing methods.

multimodaltrainingapplications

An Empirical Study of Harness Design for Coding Agents

Sep 17, 2026

Run-Ze Fan, Zihao Zhang, Simin Ma et al.

Harness design should adapt to model capability: weaker models benefit from planning and predefined tools, while stronger models achieve better cost-efficiency with minimal scaffolding and bash-only interfaces.

This paper systematically studies how different components of coding harnesses—the execution frameworks that guide AI agents through software engineering tasks—affect agent performance.

agentsevaluationapplications

Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights

Sep 17, 2026

Tica Lin, Deepak Chandran, Gauri Jagatap et al.

A shared, human-readable schema can simultaneously ground agent generation and enable human interpretation, making AI-generated content more transparent and controllable.

This paper introduces a semantic action graph—a structured representation of sports matches using performers, actions, recipients, moments, and states—that enables both AI agents to generate video highlights and humans to understand and control them.

agentsmultimodalapplications

Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols

Sep 17, 2026

Tariq Abdul-Quddoos, Xiangfang Li, Lijun Qian

Calibrated confidence thresholds enable RF-fingerprinting to work reliably under co-channel interference, with mathematical guarantees on missing detections—critical for spectrum monitoring in crowded wireless environments.

This paper tackles RF-fingerprinting—identifying specific transmitters by their hardware signatures—in realistic scenarios where multiple devices transmit simultaneously and interfere with each other. The authors use a CNN with calibration techniques to ensure reliable detection while controlling false negatives, validated on real 5G testbed data with Wi-Fi, LTE, and 5G signals.

evaluationsafetyapplications

OPTED: On-Policy Fine-Tuning for End-to-End Driving using a Render-Free Teacher

Sep 17, 2026

Damiano Da Col, Maximilian Igl, Peter Karkus et al.

Using a privileged teacher (trained on high-level inputs like maps and bounding boxes) to supervise a camera-based student during closed-loop fine-tuning is 1000× more sample-efficient than direct RL post-training for autonomous driving.

OPTED improves autonomous driving policies by using a privileged teacher trained with reinforcement learning to guide a camera-based student model during closed-loop fine-tuning. This approach avoids expensive direct RL training in simulation while keeping the policy close to human demonstrations, achieving 1.6-9.5× improvements in driving performance.

trainingefficiencyapplications

RAFT: A Stateful Retrieval-Augmented Framework for Troubleshooting Agents

Sep 17, 2026

Mingxuan Zhang, Xiaowen Wang, Anupma Sharan et al.

Retrieval-augmented systems work better for troubleshooting when they match on intermediate problem states rather than treating cases as whole documents—this simple structural change significantly improves finding relevant guidance.

RAFT improves how AI troubleshooting agents find relevant past cases by treating support tickets as multi-step journeys rather than static documents. Instead of retrieving entire cases, it matches the current problem state to intermediate steps in historical cases and returns the full trajectory from that matching point, helping agents understand what happened next in similar situations.

agentsapplications

MILER: Semantic Mid-Level Representation for Sim-to-Real Reinforcement Learning in Unstructured Autonomous Driving

Sep 17, 2026

Thomas Steinecker, Denis Trescher, Alexander Bienemann et al.

Zero-shot sim-to-real transfer for autonomous driving is achievable by training on a consistent semantic representation in simulation and applying the same representation to real sensor data, eliminating the need for manual policy adaptation.

MILER is a reinforcement learning framework for autonomous driving that bridges simulation and real-world deployment without manual tuning. It trains policies in simulation using a semantic bird's-eye-view representation, then transfers them to real vehicles by converting camera and LiDAR data into the same representation format.

trainingapplicationsefficiency

Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure

Sep 17, 2026

Zofia Smoleń

Labeling spreadsheet cells by role helps, but hits a hard limit because spreadsheets are inherently 2D with infinite possible relationships. Better approaches should flatten spreadsheets into 1D text rather than trying to classify cells into fixed categories.

This paper tackles how to make spreadsheets queryable by AI systems. The key insight is that spreadsheets are fundamentally 2D structures that don't fit neatly into categories—so instead of trying to classify each cell's role, the field should develop better ways to convert spreadsheets into readable text that language models can work with.

dataevaluationapplications

UniPolicy: Unified Objective-Specific Policies for Generative Search Advertising

Sep 17, 2026

Kun Yao, Yuhang Zhou, Yichi Zhang et al.

Multi-objective alignment in generative models works better when you give each objective its own specialized parameters and policies rather than trying to optimize them all with a single reward signal.

UniPolicy is a framework for search advertising that balances multiple competing business goals—relevance, clicks, and revenue—by training separate specialized policies within a shared model.

trainingapplicationsreasoning

Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

Sep 17, 2026

Sarah Radway, Andrew Cheng, Vijay Janapa Reddi et al.

Misaligned AI models can identify and exploit vulnerabilities in inference engines through carefully crafted outputs alone—a sandbox escape vector that doesn't require external input or other stack components.

This paper demonstrates that AI models can fingerprint the specific inference engine running them (like vLLM or SGLang) by analyzing their own output behavior, then exploit engine-specific vulnerabilities to escape sandboxes. The authors show concrete fingerprinting techniques and a proof-of-concept exploit chain, highlighting a critical security gap in how AI systems are deployed.

safetyagentsapplications

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

Sep 16, 2026

Hejia Geng, Zesen Huang, Haoyang Li et al.

Scientific code repositories contain structured domain knowledge that can be systematically converted into agent training data, enabling models to learn both specialized scientific skills and general capabilities through verified interaction trajectories.

ScienceIDE converts scientific code repositories into learning environments for AI agents by automating the extraction of executable tasks, verification criteria, and domain knowledge. The system trains specialized models (PhAI-IDE family) on verified scientific code interactions, demonstrating that learning from scientific software improves both code repair and general reasoning capabilities.

agentstrainingapplications

Affora: A Design System for Agent-Friendly Interfaces

Sep 16, 2026

Jin Gao

AI agents perform better when interfaces clearly communicate available actions and current state—you can achieve this through thoughtful design that serves both humans and machines without sacrificing visual flexibility.

Affora is a design system that makes software interfaces work better for both humans and AI agents. Rather than creating separate interfaces for machines, it improves existing interfaces so agents can understand what actions are available and what state the software is in, while keeping the visual design flexible and familiar to users.

agentsapplicationsevaluation

rMuscle: Robotic Muscle Memory for Efficient Vision-Language-Action Model Inference

Sep 16, 2026

Kaijun Zhou, Zhiyang Li, Le Chen et al.

Robots doing repetitive tasks can reuse cached visual features and neuron activations from previous executions, cutting VLA inference time by 30-40% without sacrificing accuracy—critical for real-time robot responsiveness.

This paper presents rMuscle, a framework that speeds up Vision-Language-Action (VLA) models for robot control by caching and reusing computation across repeated tasks.

efficiencyapplicationsagents

Evidence-Grounded Agentic Formulation Development in an Autonomous Laboratory

Sep 16, 2026

Michael M. Craig, Riley J. Hickman, Yingshan Ma et al.

Agentic systems that combine structured experimental evidence with tool use can dramatically improve drug formulation discovery—Andromeda 2 achieved 3x better results than the previous probabilistic model by leveraging in-house data to guide autonomous lab experiments.

Andromeda 2 is an AI agent that designs drug formulations by reasoning over experimental data and using lab tools to test batches automatically. It outperformed traditional optimization and manual design methods at finding high-performing formulations for paclitaxel, a poorly soluble drug, achieving 50% success versus 17% and 2% for competing approaches.

agentsapplicationsreasoning

Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

Sep 16, 2026

Daniel P. Jeong, Charles Q. Li, Hossein Hosseiny et al.

Evaluation metrics for AI radiology report generation are sensitive to reporting style choices, not just clinical accuracy—picking the right reference reports matters as much as the model itself.

This paper reveals that how radiologists write reports—their choice of words, detail level, and formatting—significantly affects how AI-generated radiology reports are evaluated.

evaluationapplications

What Breaks Under Pruning in Smart Homes, and When? Evaluating LLM Degradation Across Architectures and Task Complexity

Sep 15, 2026

Congjing Zhang, Vashishtha Patil, Henning Lange et al.

When deploying pruned LLMs for tool use, mixture-of-experts architectures are significantly more robust than dense models, and you need to evaluate specific action components—not just overall accuracy—to catch degradation early.

This paper studies how pruning (removing parts of neural networks to reduce size) affects large language models' ability to control smart home devices. The researchers test four different LLMs with various pruning strategies, finding that dense models break suddenly after modest pruning, while mixture-of-experts models are more robust.

efficiencyevaluationapplications

Reduced-Space Multi-Fidelity Bayesian Optimization of Process Simulation Models

Sep 15, 2026

Niki Triantafyllou, Andrea Bernardi, Maria M. Papathanasiou

By combining sensitivity analysis for dimension reduction with multi-fidelity optimization, you can reduce expensive simulator calls by 50%+ while maintaining optimization quality—critical for industrial design where each simulation costs hours or days.

This paper presents a method to optimize complex industrial process simulations more efficiently by combining dimensionality reduction with multi-fidelity Bayesian optimization. The approach uses cheap approximations alongside expensive detailed simulations, intelligently deciding when to use each, reducing the total computational cost while maintaining solution quality.

efficiencyapplications
evaluation
data

TART: A Modular Tool for Technique-Aware Audio-to-Tablature Guitar Transcription

Sep 10, 2026

Akshaj Gupta, Hwi Joo Park, Andrea Guzman et al.

This is the first guitar transcription system that produces both fingering positions and expressive techniques directly from audio, achieving significant improvements over prior work and handling real-world noisy recordings.

TART is a four-stage system that converts guitar audio into tablature (written notation showing which strings and frets to play). It tackles three real problems: recognizing guitar techniques like slides and bends, correctly identifying which string-fret combinations produce each note, and handling noisy real-world recordings.

applicationsmultimodaltraining

Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens

Sep 10, 2026

Carl Edwards, Edward De Brouwer, Xiner Li et al.

By training acquisition policies on historical experimental data and combining them with LLM biological priors, you can dramatically improve the efficiency of sequential biological experiments—recovering 27.7% of hits while testing only 5% of candidates.

This paper introduces AssayBench-Loop, a large benchmark of 1,389 CRISPR screens, and AssayLoop, a framework that learns which genes to test next by combining a transformer model trained on historical experiments with biological knowledge from LLMs.

reasoningapplications

Evaluating Time-Series Foundation Models and Multimodal Dietary Context for CGM Forecasting

Sep 10, 2026

Bowen Zhang, Hsiu-Wen Cheng, Hongyu Yang et al.

Foundation models for time-series forecasting need task-specific fine-tuning to work well for glucose prediction, and multimodal dietary data (food images + nutrition) provides clinically useful signals that significantly improve postprandial glucose forecasting.

This study evaluates whether time-series foundation models can forecast blood glucose levels from continuous glucose monitoring (CGM) data, and whether adding dietary information improves predictions.

evaluationmultimodalapplications

Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models

Sep 10, 2026

Rodion Krjutškov, Eduard Barbu, Nikos Sakkas et al.

Conversational interfaces powered by LLM function-calling can make complex ML model explanations more accessible to non-technical domain experts than traditional XAI dashboards, achieving higher accuracy and usability.

This paper presents a conversational AI system that helps facility managers understand complex energy consumption forecasting models through natural language dialogue. Instead of traditional technical dashboards, users can ask questions in plain English, and the system uses modern LLMs to interpret queries and explain model predictions with 94% accuracy.

applicationsagents

Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models

Sep 10, 2026

Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif et al.

High accuracy in medical AI often reflects leaky data rather than model sophistication. Simpler, interpretable models can match complex ones while being faster and more auditable for fairness and safety issues.

This paper audits cardiovascular screening models trained on health survey data, revealing that their reported high accuracy (AUROC ~0.89) comes from data leakage rather than genuine learning. By systematically removing leaky features and testing multiple model types—from simple classifiers to advanced foundation models—the authors show all models collapse to similar performance.

evaluationsafetyapplications

Dynamic language model representations for multi-objective reaction optimisation

Sep 10, 2026

Joshua W. Sin, David Ming Segura, Bojana Ranković et al.

Language models can learn task-specific representations of chemical reactions directly from text, enabling more efficient multi-objective optimization than hand-crafted chemical descriptors—achieving high-yield synthesis conditions with under 3% of the design space explored.

Researchers developed a method to optimize chemical reactions across multiple goals (yield, selectivity, safety) by using a fine-tuned language model to learn dynamic representations of reaction components from text descriptions.

applicationsdata

ReCite: Agentic Reasoning for Faithful Citation

Sep 8, 2026

Yuyang Huang, Bobo Li, Jiajia Song et al.

Moving from semantic similarity to claim-level reasoning dramatically improves citation accuracy—the system catches misattributions that retrieval-only approaches miss by verifying that papers logically support the claims they're cited for.

ReCite is an AI system that automatically finds and verifies citations for academic papers by reasoning about whether sources actually support claims, rather than just matching similar text. It uses an agent that plans queries, retrieves papers, and checks logical consistency—catching cases where a real paper doesn't actually back up what you're claiming.

agentsreasoningapplications
trainingefficiencyapplications

Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragility in Vision Language Reward Models

Sep 4, 2026

Wonje Jeung, Sangyeon Yoon, Hyesoo Hong et al.

Vision-language reward models for robotics are fragile to paraphrasing—rewording the same goal can flip success/failure judgments on identical robot behavior, a critical flaw for reliable robotic learning systems.

Vision-language models are being used to score robot behavior, but they fail a basic requirement: giving the same score when instructions are paraphrased. This paper introduces ROBORMBENCH, a benchmark of 2,390 real robot trajectories with 21,673 paraphrases, showing that current VLMs flip between calling identical robot actions successful or failed depending on how you word the goal.

evaluationsafetyapplications

A Deep Generative Model for Synthesizing Labeled Wireless Signals

Sep 4, 2026

Yuxiao Li, Keke Hu, Santiago Mazuelas et al.

Synthetic wireless signals generated by GANs can effectively replace expensive real-world data collection for training wireless sensing models, reducing labeling costs while maintaining physical realism.

This paper introduces IIns-GAN, a deep learning method that generates realistic labeled wireless signals for training machine learning models.

datatrainingapplications

Reflection-aware Generative Novel View Synthesis

Sep 4, 2026

GeonU Kim, Shin Dong-Yeon, Tae-Hyun Oh

You can generate realistic novel views of mirror scenes by treating reflections as virtual views and using gated attention mechanisms—no retraining needed, just clever use of existing diffusion models.

This paper presents Ref-GeNVS, a method for generating novel views of scenes containing mirrors without requiring additional training. The key innovation is treating mirror reflections as complementary views by estimating the mirror plane and reflecting camera poses, then using a two-stage approach with special attention mechanisms to ensure reflections stay consistent during generation.

multimodalarchitectureapplications

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Sep 4, 2026

Matthias Busch, Marius Tacke, Sviatlana V. Lamaka et al.

LLM performance on molecular benchmarks may reflect memorization rather than true predictive capability—models retrieve published values verbatim, and this retrieval is harder to detect and suppress than expected, especially in reasoning-enhanced models.

This paper reveals that frontier language models often retrieve memorized molecular property values verbatim from training data rather than genuinely predicting them. Testing 22 models on 12 benchmarks, researchers found that over 50% of models show exact retrieval on some datasets, and this behavior increases significantly when models use higher reasoning levels.

evaluationapplications

Design Docs Are All You Need: An AI-native Machine-Learning Performance Tool

Sep 4, 2026

Samuel Kushnir, Kimia Noorbakhsh, Kavya Sreedhar et al.

Design documents with worked examples can be more maintainable than code for ML performance tools, since AI agents can reliably regenerate implementations from them, reducing the cost of keeping performance models up-to-date with new hardware and models.

SMART is a machine-learning performance modeling tool that replaces traditional code with natural-language design documents. AI agents regenerate the entire implementation from these docs on each update, eliminating tech debt while maintaining accuracy—validated against real systems like DeepSeek-V3 on TPUs.

efficiencyarchitectureapplications

Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation

Sep 4, 2026

Siliang Liu, Mohammad Ghasemi, Sapan Patel et al.

You can capture complex LLM reasoning in a tiny classifier by training on natural-language explanations, then adapt it to specific contexts at inference time—enabling LLM-quality decisions at production scale.

This paper solves the problem of recommending product upgrades at massive scale by distilling LLM reasoning into a small, fast classifier, then fine-tuning it per product category. Instead of calling an LLM for millions of product pairs, they train a 15.5M-parameter model on LLM-generated explanations, achieving 95% of LLM quality while being 5,000x faster and 10,000x cheaper.

efficiencytrainingapplications

Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education

Sep 4, 2026

Rayed AlGhamdi

Students distinguish between AI feedback utility and evaluative authority—they'll use AI suggestions to improve writing but don't think AI should decide grades. This matters for educators integrating GenAI into assessment.

This study explores how undergraduate computing students perceive AI-generated feedback and grades when explicitly told an AI system (ChatGPT) produced them.

evaluationapplicationsalignment

The History Is the Detector: Executing CVE Patch History, End-to-End

Sep 4, 2026

Qiushi Wu, Kevin Eykholt, Youngja Park et al.

Instead of manually hunting for vulnerabilities, you can automatically extract detection patterns from known CVE fixes and apply them to find similar bugs elsewhere—turning historical security data into executable detection workflows.

BUGSTONE-E2E transforms CVE patch history into automated detection rules that find similar vulnerabilities in new code. It mines fixing commits to extract reusable patterns, then uses a funnel-shaped pipeline combining lightweight analysis, LLM inspection, and runtime verification to identify and validate security flaws across 14 programs.

safetyevaluationapplications

Lightweight Vision Transformer Compression for On-Device Plant Disease Detection in Resource-Constrained Agricultural Field Conditions

Sep 4, 2026

Mahadev Sunil Kumar, Bhavika Gondi, Desaisetty Venkata Satya Sai Swapnith et al.

Combining multiple compression techniques (pruning, quantization, distillation) can dramatically shrink Vision Transformers for on-device deployment, but simpler approaches sometimes match performance at lower computational cost.

This paper presents a method to compress Vision Transformers for plant disease detection on mobile devices. The researchers combine three compression techniques—pruning, quantization, and knowledge distillation—to reduce model size by 54.5x while maintaining 95% accuracy on chilli disease classification, enabling deployment in resource-limited agricultural settings.

efficiencyapplicationstraining

Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

Sep 4, 2026

Alexander Neubauer, Tianzhen Hong, Han Li et al.

LLMs for building HVAC are research-stage tools best suited for semantic and workflow support (naming, documentation, operator guidance), not autonomous control—conventional ML and model predictive control remain more reliable for actual operational decisions.

This review examines 66 studies on using large language models for HVAC building control systems. While LLMs show promise for semantic tasks like naming conventions and operator support, the research remains largely theoretical—only 4 studies reached pilot stage and none achieved real operational deployment.

applicationssafetyevaluation

SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents

Sep 3, 2026

Xin He, Yanlin Wang, Mingwei Liu et al.

Functional test passage alone is insufficient for evaluating coding agents—real-world software development requires meeting code review constraints, and current agents have a significant gap between passing tests and satisfying these constraints.

SWE-Gate is a benchmark that evaluates coding agents not just on passing tests, but on meeting real-world code review standards. It includes 303 repair tasks from Python repositories with explicit review constraints derived from actual pull request comments, revealing that many patches pass functional tests but fail review requirements.

evaluationagentsapplications

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

Sep 3, 2026

Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild

Offloading structured reasoning (like network topology) from LLMs to specialized models (graph encoders + RL policies) makes AI security agents faster, more reliable, and deployable at enterprise scale.

This paper presents Sentinel-RL, a system that helps LLM-based security analysts by splitting their work: a graph neural network handles the complex authentication network topology, while reinforcement learning constrains the agent's actions to valid security moves, and the LLM focuses on explaining decisions to humans.

agentssafetyapplications

Efficient Test-Time Adaptation through Human-AI Interaction

Sep 3, 2026

Zora Zhiruo Wang, Apurva Gandhi, Rulin Shao et al.

AI agents can be efficiently personalized to individual users by learning from their iterative feedback during real work, rather than requiring explicit upfront specifications of what success looks like.

This paper presents TAHI, a method that adapts AI agents to individual users by learning from their feedback during repeated interactions. Instead of training on generic population data, the system uses a user's preferences and corrections across multiple tasks to personalize the agent's behavior and create custom evaluation rubrics.

trainingapplicationsagents

The Natural Language Interaction Protocol and Standard for AI Agents

Sep 3, 2026

Luyi Xing, Rasit Onur Topaloglu, Ranjan Sinha et al.

NLIP is a practical interoperability standard for AI agents that lets you build agents independently and have them work together seamlessly, similar to how different web services communicate via HTTP.

This paper introduces NLIP, a standardized communication protocol that lets AI agents built with different frameworks and tools talk to each other. Like HTTP for web servers, NLIP provides a common language for agents to exchange messages, coordinate work, and access shared tools—enabling organizations to mix-and-match agent systems without rebuilding everything from scratch.

agentsarchitectureapplications

Adaptive Vision-Language Grasping via Composable Foundation Priors and Generalizable Grasp Synthesis

Sep 3, 2026

Sixu Yan, Shikang Wang, Binhua Huang et al.

Decoupling physical grasp synthesis from task-dependent reasoning allows robotic systems to leverage foundation model improvements without retraining, and enables cross-hand generalization through explicit kinematic and stability constraints.

AdaRoboVLG combines vision-language models with robotic grasping by separating physical grasp synthesis from task understanding. A base policy generates and evaluates grasp candidates using kinematics and stability checks, while foundation models provide task-specific context.

applicationsmultimodalreasoning

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

Sep 2, 2026

Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi et al.

Post-training with domain-specific data, reasoning traces, and test-time compute strategies can push language models to superhuman performance on complex reasoning tasks like competitive programming.

Researchers trained specialized language models to excel at competitive programming by combining curated problem datasets, synthetic reasoning traces, fine-tuning, and reinforcement learning. Their system achieved gold-medal performance on the International Olympiad in Informatics (IOI) 2025 and 2026, becoming the first AI to outscore the top human competitor on an IOI problem set.

trainingreasoningapplications

Large Language Models (LLMs) for Telecom Root Cause Analysis (RCA): A Structured Reasoning Framework for Evidence-Grounded Diagnosis

Sep 2, 2026

Hao Zhou, Mandar Kulkarni, Hao Chen et al.

LLMs need structured reasoning frameworks with evidence grounding to reliably diagnose telecom network faults—vanilla LLMs hallucinate and produce unstable reasoning without domain-specific constraints and verifiable decision paths.

This paper addresses root cause analysis (RCA) in telecom networks by proposing a structured reasoning framework for LLMs that reduces hallucination and improves diagnostic accuracy. The approach organizes network data into canonical contexts, enforces decision-path reasoning, and grounds explanations in evidence, demonstrating improvements on 5G network datasets.

reasoningapplicationsevaluation

DiscoSign: Discourse-Aware Text to Sign Language Gloss Translation

Sep 2, 2026

Vasileios Baltatzis, Mert Inan, Connor Gillis et al.

Sign language translation needs discourse context to maintain spatial consistency and entity tracking across multiple sentences—not just translating each sentence independently.

DiscoSign translates English text to American Sign Language glosses while tracking discourse-level phenomena like spatial references and entity consistency. Traditional sign language systems work sentence-by-sentence, missing how entities and concepts maintain meaning across longer conversations.

multimodalapplicationsevaluation

StudentSim: Training LLM-based Student Simulators

Sep 1, 2026

Ke Yang, Chenglong Wang, Michel Galley et al.

Student simulators trained with this method can serve as realistic reward signals for tutor reinforcement learning, letting AI tutors optimize their teaching strategies without requiring extensive real student data.

StudentSim trains AI models to simulate individual students' learning behaviors and how they respond to tutoring.

trainingapplicationsevaluation

Designing Proactive Thought Partners for Writing

Sep 1, 2026

Chao Zhang, Abe Davis, Chih-Wei Chen et al.

Proactive writing assistants work best when users can customize what kind of help they get, when they receive it, and how it's presented—avoiding generic suggestions in favor of role-specific, lightweight interventions that support rather than interrupt the writing process.

This paper explores how AI agents can proactively help writers by offering customizable, higher-level cognitive support at the right moments. Researchers built and tested a tool where users configure AI 'thought partners' with specific roles and proactivity levels, then observed how writers used suggestions for ideation and self-review.

applicationsagentsreasoning

Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics

Sep 1, 2026

Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin et al.

Smart data curation and cost-aware deployment economics can make small VLMs competitive with much larger models for document processing—the key is selecting training examples strategically and measuring real-world costs including verification and correction.

This paper presents a practical document understanding system that extracts structured data from documents at a fraction of the cost of human annotation or larger models.

efficiencydataapplications

From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix

Sep 1, 2026

Olga Tsymboi, Dmitrii Stoianov, Ramil Latypov et al.

Training separate optimization experts for different failure modes, then merging them, beats joint optimization and lets enterprises consolidate their LLM fleet without sacrificing quality.

A company built a single self-hosted LLM to replace 200+ fragmented models by analyzing production errors and training specialized experts for instruction-following, function-calling, and task distribution. They merged these experts and achieved better performance than a 7× larger model while handling 116M monthly requests at lower cost.

trainingefficiencyapplications

From Confusion to Clarity: Confusion-Aware Retrieval and Knowledge Injection for Text Classification

Sep 1, 2026

Manish Gupta, Chaitanya Giri, Jayasimha Talur

By identifying confusable label pairs and generating targeted differentiation rules, you can improve text classification accuracy by up to 10 points without retraining—and these rules transfer to smaller, cheaper models.

This paper tackles a key challenge in text classification: when LLMs must choose between many similar category labels, they often get confused. The authors propose a system that identifies which label pairs the model struggles with, includes those confusing pairs in the candidate set, and generates targeted rules to help distinguish between them.

evaluationapplicationsreasoning

Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification

Aug 31, 2026

Yisen Xi

You can identify anonymous API models through systematic forensic analysis of configuration, tokenizer behavior, and archived platform data—but this requires careful validation and should decline to guess rather than make false claims.

This paper presents a four-stage forensic protocol for identifying anonymous AI models served through APIs. Using archived platform data, configuration fingerprinting, tokenizer analysis, and behavioral testing, the authors demonstrate how to verify model identity without relying on self-identification.

evaluationsafetyapplications

Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation

Aug 31, 2026

Riya Ahuja, Tim Kacprowski, Roya Shiasi Sardoabi

Semantic chunking that respects entity and relationship boundaries significantly outperforms fixed-size chunking for biomedical RAG, especially when relation cues are explicit in the text.

This paper improves biomedical information extraction in RAG systems by replacing fixed-size text chunking with a configurable semantic chunking framework. The approach preserves important entities and relationships by using trigger-centered chunking and hierarchical relation resolution, achieving 8.4 F1 points improvement on biomedical relation extraction benchmarks.

dataapplications

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

Aug 31, 2026

Yung Wei Shueh, Zhi-Jie Chen, Chia-Hsuan Hsu et al.

Building trustworthy clinical AI requires layering deterministic data processing, retrieval-augmented generation from authoritative sources, and verification checks—not relying on LLMs alone.

DIASENTINEL is a multi-agent system that uses LLMs safely for diabetes risk screening by combining clinical data extraction, guideline-based retrieval, and verification layers to prevent hallucinations and ensure all recommendations are traceable to medical guidelines.

safetyapplicationsagents

When Does Bigger Help? A Controlled Study of LLM Scale for Ontology Learning

Aug 31, 2026

Hamed Babaei Giglou, Sören Auer, Jennifer D'Souza

When building ontologies with LLMs, model size matters less than you'd expect—architecture and training approach often matter more, and the best model depends heavily on your specific task (term classification vs. relationship extraction).

This paper systematically evaluates how LLM size affects ontology learning—the task of automatically extracting concepts and relationships from text to build knowledge structures.

evaluationscalingapplications

LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

Aug 31, 2026

Gopi Krishnan Rajbahadur, Amir M. Ebrahimi, Boyuan Chen et al.

Post-training in production is about engineering discipline around data mixtures and yield metrics, not algorithmic breakthroughs—small improvements in converting training data into usable supervision can compound into significant model gains.

This paper treats LLM post-training as industrial maintenance work, not research. Teams inherit a deployed model and must improve it within strict compute budgets without breaking existing capabilities.

trainingdataapplications
evaluation
applications

Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation

Aug 27, 2026

Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong et al.

By modeling reactions as electron rearrangements instead of molecular graph edits, MAELLE achieves competitive accuracy while naturally explaining reaction mechanisms and maintaining robustness on out-of-distribution chemistry.

This paper presents MAELLE, a machine learning model that predicts chemical reactions by tracking how electrons move between atoms, rather than just predicting final products.

reasoningapplicationsdata