ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2582 papers29 this month12 topics
AllTraining 48Reasoning 39Efficiency 35Evaluation 34Agents 25Applications 19Multimodal 15Architecture 14Safety 13Data 9Alignment 3scaling 3

Oct 5 – Oct 11(9)

Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective

Oct 6, 2026

Kevin Zhang, Stephen Bates

Conformal prediction set sizes have a rigorous information-theoretic interpretation: they quantify information gain in a way that's mathematically sandwiched between generalized entropy measures and obeys data processing inequalities.

This paper establishes a theoretical connection between conformal prediction (a method for uncertainty quantification) and information theory. The authors show that the size of prediction sets from conformal methods can be interpreted as a measure of information gain, providing mathematical justification for using set size as an uncertainty metric.

evaluationsafetyreasoning

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

Oct 6, 2026

Ankit Sonthalia, Haritz Puerto, Alexander Rubinstein et al.

LLM agents struggle to convert their general capabilities into cost-efficient task-specific solutions, but when they succeed, the savings are dramatic—suggesting bottling is a valuable but underdeveloped capability worth improving.

This paper introduces BOTTLED, a benchmark testing whether LLM agents can autonomously create cheaper, task-specific solutions from their general capabilities. Agents receive unlabeled workloads with fixed budgets and must decide their own approach—like training small models or writing programs.

Sep 28 – Oct 4(31)

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Oct 2, 2026

Ruihong Shen, Žiga Kovačič, Peter Kulits et al.

Strong vision models fail at understanding dynamic scenes through code generation—the gap between static and dynamic reconstruction is a major frontier for AI agents.

4DCodeBench is a benchmark that tests AI agents on reconstructing dynamic 3D scenes from video by writing graphics code. Agents must understand physics, deformation, and fluid dynamics to generate executable programs that recreate what they see. The benchmark reveals that current models struggle with complex motion even when they're good at static scene reconstruction.

evaluationreasoningagents

What Should World Models Forget? Stratified Retention for Continual Adaptation

Oct 2, 2026

Nishit Anand, Ramani Duraiswami, Dinesh Manocha

World models need stratified forgetting strategies that preserve physical invariants while quickly adapting to environmental changes—standard continual learning metrics fail to capture this distinction and incorrectly reward frozen models.

This paper addresses a fundamental problem in continual learning for world models: knowing what to forget. Unlike traditional learning where correct labels stay correct, world models operate in changing environments where outdated knowledge must be discarded.

Sep 21 – Sep 27(29)

First-Order Stationarity of Reverse Diffusions

Sep 25, 2026

Zhifeng Chen, Chenyang Jiang, Yazhen Wang

Diffusion models have provable convergence guarantees similar to optimization algorithms—reverse diffusions contract divergence exponentially fast, and discrete samplers achieve measurable stationarity bounds that don't depend on data convexity.

This paper connects optimization theory to diffusion models by proving that reverse-time diffusion processes contract Fisher divergence at exponential rates under strong convexity conditions. The authors also establish first-order stationarity bounds for practical discrete samplers, showing how optimization guarantees translate to sampling quality without requiring global convexity.

trainingevaluationreasoning

Statistical attribute alignment for black-box generative AI via output post-processing

Sep 25, 2026

Kevin Jiang, Morgane Austern, Edgar Dobriban et al.

You can align generative AI outputs to target distributions by intelligently filtering multiple model queries—no model retraining needed—and this approach is provably optimal for large batches of outputs.

This paper addresses how to align AI-generated outputs with user-specified attribute distributions through post-processing, without modifying the model itself. The authors develop algorithms that select outputs from multiple queries to a generative model, ensuring attributes like gender or age match target distributions.

Sep 14 – Sep 20(31)

Cross-sector generalization of accident-process role classification in occupational accident narratives

Sep 18, 2026

Aho Yapi, Pierre Latouche, Arnaud Guillin et al.

Fine-tuned language models can generalize accident report classification across different industries without retraining, enabling scalable occupational safety analysis across sectors.

This paper develops an automated system to classify key information in occupational accident reports (work situations, unsafe conditions, events, consequences) and tests whether models trained on construction-sector narratives can work across different industries like metallurgy and chemistry.

trainingevaluationapplications

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

Sep 18, 2026

Renkai Ma, Ruyuan Wan, Xuan Lu et al.

Building AI agents that users trust requires focusing on operating conditions—cost, oversight, and access controls—not just task performance. Users care deeply about being able to supervise and review agent actions.

This study analyzed 73,000+ Reddit posts about using OpenClaw (an AI agent tool) to understand what values matter to users beyond just task completion.

agentsefficiencyevaluation

The Missing Minimal Pair: Stereotype Evaluation in LLMs

Oct 6, 2026

Nataliya Stepanova, Ivan Titov, Emily Allaway et al.

Single stereotype sentence pairs are unreliable for measuring LLM bias; use dual minimal pairs and mutual information-based metrics instead for consistent, language-agnostic bias evaluation.

This paper identifies a flaw in how bias is typically measured in language models: comparing just two sentences about stereotypes can give contradictory results depending on how you rephrase them. The authors propose a better approach using dual comparisons and introduce metrics based on mutual information to more reliably measure stereotype bias across languages and models.

evaluationsafety

A Systematic Study of Semantic ID Spaces for Generative Information Retrieval

Oct 6, 2026

Alexia Allal, Hicham Randrianarivo, Sylvain Lamprier

DocID design significantly impacts generative retrieval performance—the paper provides a systematic framework and training-free metrics to evaluate and optimize DocID structures, enabling faster iteration than traditional downstream evaluation.

This paper systematically studies how to design effective document identifiers (DocIDs) for generative information retrieval systems.

evaluationefficiency

One Figure, Every Canvas: Editable Flowchart Relayout via Agentic Pipeline

Oct 5, 2026

Shih-Chen Tseng, Chih-Hsuan Chen, Ryan Yang et al.

Agentic systems with explicit constraint checking and visual critics can reliably preserve structural integrity in document layout tasks—achieving 68.6% fidelity versus 11-41% for prior methods—by factoring the problem into specialized stages rather than end-to-end generation.

This paper tackles the problem of automatically adapting flowchart diagrams to different aspect ratios (like fitting a pipeline figure into a paper column, slide, or social media format) while preserving all connections and content.

agentsapplicationsevaluation

BiasFlow: Geometric Monitoring and Backbone Regularization for Spurious Feature Reliance

Oct 5, 2026

Haojin Deng, Zhiping Lin, Yimin Yang

Monitoring centroid geometry during training can help detect and reduce spurious feature reliance, but attribute information remains partially recoverable—suggesting regularization alone isn't sufficient for complete bias removal.

BiasFlow is a monitoring toolkit that tracks how neural network backbones rely on spurious features (like gender in face recognition) through geometric analysis of feature centroids.

safetyevaluationtraining

PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

Oct 5, 2026

Yaohui Zhang, Binxu Li, Haoyi Duan et al.

Current multimodal models can approximate values from scientific figures but struggle with precision; providing source data instead of figures dramatically improves accuracy (90% to 97.4%) while reducing computational cost, suggesting a practical path for scientific data extraction.

PlotGround is a benchmark for evaluating how well AI models can extract numerical values from scientific figures.

evaluationmultimodaldata

TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts

Oct 5, 2026

Oliver Jaffe, Dane Sherburn

AI models are becoming exponentially better at the experimental research process itself (not just raw capability), with frontier models now reaching expert-level results using 2.3x less compute—a skill that could significantly accelerate AI R&D timelines.

TasteVal is a benchmark measuring how efficiently AI models design and conduct experiments to solve research problems. Rather than evaluating raw problem-solving ability, it measures 'experimental taste'—the skill to iteratively design good experiments and interpret results—by comparing how much compute a model needs versus human experts to reach the same performance level.

evaluationreasoningagents

Deep Learning for Sleep Heart Rate Estimation from Accelerometers: Toward Population-Scale Cardiac Insight Without Optical Sensors

Oct 5, 2026

Tanbin Islam Rohan, Pranjol Sen Gupta, Tanusree Debi et al.

You can estimate sleep heart rate from accelerometer motion signals using deep learning, trading off some accuracy for broader coverage—useful for extracting cardiac insights from existing wearable data without optical sensors.

This paper presents SeqSmoother, a transformer-based model that estimates heart rate during sleep using only wrist accelerometer data, without requiring optical sensors.

trainingevaluationapplications
trainingevaluationreasoning

RNADyn: A Benchmark for Generating and Understanding RNA Dynamics

Oct 2, 2026

Yiming Huang, Lennart Bastian, Hanqun Cao et al.

A unified deep learning approach can both generate realistic RNA dynamics trajectories and predict dynamics fingerprints from static structures, bridging two previously separate tasks and improving physical accuracy through explicit physical constraints.

This paper introduces RNADynBench, a large-scale benchmark of 2,585 RNA molecular dynamics simulations, and RNADynNet, a unified model that generates realistic RNA trajectories and extracts dynamics information from single structures.

dataarchitectureevaluation

Forecasting from Counterfactual Simulator Rollouts: A Sim2Real Evaluation

Oct 2, 2026

Angel Wang, Dominique Perrault-Joncas, Alvaro Maggiar et al.

Simulator-generated counterfactual rollouts can effectively bootstrap forecasting models for new policies before real deployment data exists, and these models improve further with minimal real-world calibration.

When deploying a new decision policy, prediction models face a cold-start problem because historical data reflects old policies, not the new one. This paper uses simulation to generate counterfactual training data by rolling out the new policy in a simulator, then tests whether models trained on simulated data transfer to real-world inventory control.

applicationsevaluation

Revisiting Input Time-frequency Representations in Multi-pitch Estimation for Vocal Ensembles

Oct 2, 2026

Junyoung Koh, Hao-Wen Dong

Linear STFT outperforms the more complex HCQT for vocal ensemble pitch estimation while reducing computational cost—simpler input representations can be more effective when properly designed.

This paper challenges the conventional use of harmonic constant-Q transform (HCQT) for multi-pitch estimation in vocal ensembles by showing that simpler linear STFT representations actually perform better while being much faster to compute.

evaluation

MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search

Oct 2, 2026

Sean Culatana, Shang-En Huang, Kang Li

MRVQ enables one quantized index to serve all dimension-rate combinations by using truncatable residual quantization, reducing memory overhead by 17.8-22x compared to training separate indices, though with modest quality trade-offs.

This paper introduces MRVQ, a quantization method that compresses embeddings for vector search while supporting flexible trade-offs between embedding dimension and compression rate. A single index can be truncated in two ways—dropping quantization stages or embedding coordinates—to adapt to different memory and latency constraints without storing multiple separate indices.

efficiencyevaluation

Do Large Language Models Know Colombian Law? A Reliability Benchmark for the Colombian Legal System

Oct 2, 2026

Rubén Manrique, Michelle Castellanos, Jorge Morales et al.

LLMs can sound authoritative about law they don't actually know; current models need expert oversight and source grounding for real legal work, especially outside the US where training data is sparse.

This paper evaluates how well large language models understand Colombian law by testing 15 models on 1,042 expert-validated questions covering ten legal areas. While models score well on multiple-choice questions (up to 90.5%), their free-text legal answers are rarely correct (max 45%), and they often sound confident while being wrong—a dangerous combination for non-experts relying on legal AI.

evaluationsafetyapplications

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Oct 1, 2026

Pengfei Li, Naufal Suryanto, Sicheng Zhang et al.

Current open-weight LLMs struggle with precise cybersecurity tool use (max 42% accuracy), but fine-tuning with verifiable rewards from this benchmark can make smaller models competitive with much larger ones.

KaliBench is a benchmark for evaluating how well language models can translate security analyst requests into executable commands for Kali Linux tools. It includes 8,504 query-command pairs across 1,642 tools and provides a verification system that checks both whether commands are syntactically correct and whether they actually run successfully, without needing to execute them during training.

evaluationagentssafety

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Oct 1, 2026

Sohyeon Kim, Yoonho Lee, Bo Liu et al.

Even advanced AI agents fail at retrieving papers that inspired real research (max 0.51 recall), revealing a critical gap in how models search scientific literature—this task requires something beyond current retrieval and reasoning approaches.

ScholarCatalyst is a benchmark dataset where 184 computer science researchers labeled which prior papers inspired their completed projects. The benchmark tests whether AI systems can retrieve these influential papers given only an initial research question and literature available at project start.

evaluationreasoningagents

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

Oct 1, 2026

Shuo Xing, Zilin Dai, Chengyuan Qian et al.

Mathematical reasoning in LLMs isn't a single skill but four distinct capabilities; focusing training on the 'Discovery' bottleneck (finding the right solution strategy) is more effective than generic math training.

This paper diagnoses why LLMs struggle with math by breaking down mathematical reasoning into four components (Discovery, Generation, Digestion, Execution) and shows that Discovery—finding the right approach—is the main bottleneck. The authors then propose a training method that uses these insights to improve math performance across different model sizes.

reasoningtrainingevaluation

Every Ablation Is a Dose: Counterweights and the Semblance of Self-Repair

Oct 1, 2026

Areeb Ahmad, Pratinav Seth, Vinay Kumar Sankarapu

Self-repair in language models isn't adaptive compensation—it's pre-existing counterweights responding predictably to ablation. You can predict how a component will respond to intervention from its fixed weights alone.

When you disable a component in a language model, other parts often seem to compensate—a phenomenon called 'self-repair.' This paper shows it's not actually repair: it's pre-existing counterweights doing their normal job.

evaluation

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

Oct 1, 2026

Juan S. Santillana

Don't trust tool-use benchmarks for small models—use verbatim-reproduction checks and token-probability probes to verify genuine capability before claiming tool use works.

Small language models can appear to use tools correctly on standard benchmarks while actually just memorizing training examples. This paper shows how keyword-matching tests miss real failures, proposes cheap diagnostic checks to catch false positives, and demonstrates a targeted fix that repairs tool-use ability in a 1.1B parameter model using minimal compute.

evaluationefficiency

MIRTO: a registration-gated, multiverse-tested evaluation protocol for unsupervised anomaly segmentation in brain MRI

Oct 1, 2026

Negin Kafee Hernashki, Soumick Chatterjee

Anomaly detection benchmarks hide critical methodological choices; MIRTO shows that reported performance differences between methods often reflect evaluation setup rather than actual capability differences, and that registration alignment and threshold selection are major hidden sources of variance.

MIRTO is an evaluation protocol that makes explicit the hidden choices in unsupervised brain MRI anomaly detection—how maps are aligned, thresholds set, and metrics calculated.

evaluationsafety

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Oct 1, 2026

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing et al.

Current frontier AI models struggle with real enterprise data work: the best model scores 95+ on only 35% of tasks, revealing a major gap between text-to-SQL benchmarks and actual data agent capabilities needed for production systems.

Argo-Bench is an evaluation framework with 210 realistic data science tasks that test AI agents on enterprise-scale workflows.

evaluationagentsdata

A Comparative Explainability Framework for DeBERTa-v3 in Zero-Shot Medical Abstract Classification

Oct 1, 2026

Javier Diaz Esteban-Herreros, David Muñoz-Valero, Raquel Martínez-España et al.

When auditing AI models for medical text classification, use multiple explanation methods together and check their agreement—single methods can be misleading, especially when the model is uncertain about its prediction.

This paper evaluates how well different explanation methods agree when analyzing a medical text classifier (DeBERTa-v3). Using five different explanation techniques on medical abstracts, researchers found that explanations are most reliable when the model is confident, but diverge significantly when the model is uncertain.

evaluationsafety

Are We Recovering Mechanisms? Objective-Level Recovery Gaps in Mechanistic Interpretability

Oct 1, 2026

Chuqin Geng, Li Zhang, Haolin Ye et al.

Standard circuit evaluation metrics can systematically prefer incorrect mechanisms, making better discovery algorithms insufficient without fixing the evaluation objective itself.

This paper reveals a critical flaw in how mechanistic interpretability evaluates circuit discovery: the standard faithfulness metrics can prefer worse circuits that merely reproduce model behavior without actually capturing the underlying mechanisms.

evaluationsafety

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

Oct 1, 2026

Kyochul Jang, Seohyeon Park, Ohchul Kwon et al.

Humanoid robots struggle with tool selection and coordinating manipulation with movement—even state-of-the-art models like GR00T show reduced accuracy on unseen tools and can execute tasks despite receiving unrelated instructions.

This paper introduces HumanoidToolBench, a benchmark for evaluating humanoid robots on tool use tasks that require selecting appropriate tools and coordinating manipulation with locomotion. The benchmark includes 3,100 demonstrations and tests seven policies, revealing significant gaps between tool selection and successful task completion.

evaluationagents

Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling

Oct 1, 2026

Daniel McBride, Pratik Khandagale, Cristina Garcia-Cardona et al.

Popular diffusion-based sampling methods have a fundamental limitation for multimodal distributions—they require exponentially long times to transition between well-separated modes, making theoretical convergence guarantees misleading about practical efficiency.

This paper challenges the claim that Wasserstein gradient flows and forward-only diffusion can efficiently sample complex multimodal distributions. Using tools from statistical physics, the authors prove these methods suffer from exponentially slow mixing times when modes are well-separated, regardless of theoretical convergence guarantees.

evaluationscaling

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

Oct 1, 2026

Arman Behnam, Binghui Wang

Current memory systems can't measure the value of memories that are never retrieved.

Memory-augmented language models struggle to identify which memories are actually useful because some memories are never retrieved, making their value impossible to measure.

reasoningevaluationtraining

AI Emulation of Stochastic Sudden Stratospheric Warming with Interpretable Latent Structure

Oct 1, 2026

C. Daniel Boscu, Daniel Hernandez, Fabio Alvarez Ventura et al.

Deep learning models can be designed to both predict rare, extreme events accurately and produce interpretable internal representations that align with real physical phenomena, enabling better understanding of how AI makes decisions about complex systems.

Researchers built an AI model to predict rare weather events called Sudden Stratospheric Warming by learning from a simplified atmospheric model.

reasoningevaluationapplications

Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis

Sep 30, 2026

Tian Xia, Minghao Liu, Yiqing Liang et al.

For imbalanced clinical tasks, optimizing prompts for ranking metrics (AUROC) instead of accuracy can dramatically improve model performance—up to 16 percentage points—because accuracy-based optimization fails when one class dominates the data.

This paper addresses class imbalance in clinical diagnosis by optimizing multimodal language models for AUROC instead of accuracy. The authors introduce Ranking-PE, a prompt optimization method that evaluates candidate prompts based on how well they rank positive cases above negative cases, rather than raw correctness.

evaluationmultimodalapplications

Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text

Sep 30, 2026

Dulhan Jayalath, Oiwi Parker Jones

Brain-to-text decoders can accidentally learn from word timing patterns instead of brain activity; removing this shortcut with independent window processing makes the task genuinely harder but enables better learning from actual neural signals.

This paper reveals that a major brain-to-text decoding method was exploiting timing shortcuts from word duration patterns rather than learning from actual brain signals. By processing brain windows independently instead of jointly, the authors eliminate this shortcut and achieve better performance (36.6% word error rate) using simpler methods like prediction aggregation and language model priors.

evaluationreasoningapplications

ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

Sep 30, 2026

Xinghao Chen, Xiangbo Gao, Jiongze Yu et al.

Video text editing requires balancing three competing goals: correct text, smooth motion, and unchanged background. This benchmark and dataset help measure those trade-offs and establish baselines for the community.

ViTeX-Bench is a benchmark for video scene text editing—replacing text on surfaces like signs and labels while keeping the rest of the video unchanged. It includes 387 real videos, evaluation metrics for text accuracy and visual quality, and a baseline editor that achieves strong results. This addresses a gap where video text editing lags behind image editing.

evaluationmultimodalapplications

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

Sep 30, 2026

Ziyan Jiang, Jingbo Yang, Jiabao Ji et al.

Multimodal agents struggle to couple exploration and visual reasoning in 3D worlds—they can see anomalies or navigate, but struggle to do both together effectively, suggesting a fundamental gap in how these systems integrate action and perception.

WorldAuditBench is a benchmark for testing how AI agents find problems in 3D virtual worlds—like floating objects or walls you can walk through. It evaluates multimodal AI systems (vision-language models and vision-language-action models) on 213 anomaly detection tasks across 13 environments, measuring how well agents can explore systematically and visually identify issues.

evaluationmultimodalagents

Cogentic: Multi-Agent Orchestration for Automated Proof Discovery

Sep 30, 2026

Yang Cai, Vineet Gupta, Yanchen Jiang et al.

Multi-agent orchestration with verification loops can solve research-level math problems that single-shot generation cannot—by exploring multiple directions, catching errors through adversarial checking, and retaining progress across long reasoning horizons.

Cogentic is a multi-agent system that orchestrates teams of AI provers to tackle open research problems in mathematics and theoretical computer science. Instead of relying on single attempts, it uses an iterative loop where specialized agents explore different proof directions, verify results against each other, and build on confirmed findings stored in a persistent ledger.

agentsreasoningevaluation

Cropland PAtteRNS: Parallel Dimensional Attention Networks and Attention to Dataset Disparity for Crop Segmentation in Satellite Imagery Time Series Data

Sep 29, 2026

Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini

Factorizing attention across different data dimensions improves crop segmentation, but dataset construction choices (tile size, class definitions) have outsized impact on results and must be standardized for meaningful model comparisons.

This paper introduces PAtteRNS, a transformer-convolutional model for crop segmentation in satellite imagery that separately applies self-attention to temporal, spectral, and spatial dimensions. The authors also highlight critical dataset issues—flawed class groupings and incompatible tile-size variants—that undermine fair model comparison and suggest standardization is needed.

architecturemultimodalevaluation

A Spectral Theory of Distortion in LLM Graph Reconstruction: Sharp Bounds and Empirical Characterization

Sep 29, 2026

Jianru Shen

When evaluating LLM graph reconstruction, a single distance metric masks whether errors come from adding edges, removing edges, or both—the paper provides a mathematical framework to detect mixed editing and reveals different models have fundamentally different failure modes.

This paper analyzes how language models reconstruct graphs, proving mathematical bounds on the Wasserstein distance between original and reconstructed graph spectra. The bounds reveal whether a model only adds edges, only removes them, or does both—information hidden by standard distance metrics.

evaluationreasoning

LongHarness Bench: Stress-Testing Language Model Harnesses for Long-Context Reasoning

Sep 29, 2026

Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen et al.

Long-context evaluation needs to measure both accuracy and computational efficiency—the same model can be dramatically more or less efficient depending on its processing strategy, and current benchmarks don't capture this variation.

This paper introduces LongHarness Bench, a benchmark for evaluating how well language models handle long documents using different processing strategies. The benchmark tests both accuracy and efficiency, requiring models to find relevant information across scattered context and reason strategically rather than reading everything.

evaluationefficiency

PDMD: Projected Distribution Matching Distillation for Video Diffusion Models

Sep 28, 2026

Zimo Wang, Junkun Yuan, Angtian Wang et al.

Critic error accumulation is a fundamental bottleneck in distilling video diffusion models; filtering it via projection dramatically improves sample quality without architectural changes or extra computation.

This paper improves video diffusion model distillation by fixing a key problem: critic errors that accumulate during training and degrade sample quality. PDMD uses a simple mathematical projection to filter out these errors while preserving useful learning signals, achieving better video quality with fewer computational steps—all with just a one-line code change.

efficiencytrainingevaluation

TokenCast: Forecasting Token Consumption During LLM Agent Execution

Sep 28, 2026

Chaoqian Ouyang, Ling Yue, Libin Zheng et al.

Token consumption in agentic LLM workflows is unpredictable and can vary 10x+ per task—TokenCast forecasts it accurately by tracking execution segments and context growth, improving budget planning by 14.5% on average.

TokenCast predicts how many tokens an LLM agent will consume during task execution, which varies wildly across runs due to tool use and growing context. It learns cost patterns for each execution step and updates predictions as the agent runs, enabling better budget control without extra LLM calls.

agentsefficiencyevaluation

FinAutoRubric: Expert-Guided Automatic Rubric Generation for Evaluating Financial Research Agents

Sep 28, 2026

Hoyoung Lee, Suyeol Yun, Jack Haverty et al.

Automating rubric generation with expert oversight lets you evaluate complex financial AI systems at scale without manually writing new rubrics for each task, while maintaining quality comparable to human expert grading.

FinAutoRubric automates the creation of evaluation rubrics for financial research agents by combining expert guidance with AI-generated, task-specific criteria.

evaluationapplicationsagents
evaluationsafetyapplications

Compact Documentation for Coding Agents: A Benchmark, an Optimizer, and Why It Does Not Transfer

Sep 25, 2026

Md Shohel Arman, Igor Molybog

High-quality code documentation doesn't improve AI agents' ability to fix real bugs, suggesting that documentation quality and real-world problem-solving are decoupled—a finding that challenges assumptions about documentation's utility for coding agents.

This paper investigates whether better code documentation helps AI agents fix bugs in real repositories. The authors create a benchmark to measure documentation quality (based on whether code can be regenerated from descriptions) and an optimizer to improve it.

evaluationagentsapplications

Uncertainty and Explainability in Deep Rough Volatility: A Neural Information-Theoretic Posterior Approach

Sep 25, 2026

Damiano Brigo, Raphaël Huser, Dan Leonte

Neural calibration of financial models should quantify posterior parameter uncertainty and propagate it through pricing models—point estimates alone can be materially unreliable for exotic derivatives, and information-theoretic explainability reveals which market data regions constrain which pa...

This paper develops a neural framework for calibrating rough Heston volatility models that captures parameter uncertainty rather than just point estimates. It combines simulation-based inference with neural surrogates to produce uncertainty-aware price intervals for exotic options, and introduces an explainability method to identify which market observations drive parameter learning.

evaluationapplications

Two Conformal Constructions for Adaptive Within-Document AI-Text Screening

Sep 25, 2026

Marco Mandap, Jerahmeel Hipolito, Arcel Galvez et al.

You can build AI-text detectors that adaptively choose what to inspect and when to stop, while mathematically guaranteeing false-alert control without needing to split your error budget across all possible inspection paths.

This paper develops statistical methods to control false alerts when screening documents for AI-generated text. The authors propose two conformal inference approaches that allow adaptive inspection—selecting which parts of documents to check and when to stop—while guaranteeing that false-alert rates stay below a target threshold, even when checking multiple detection strategies.

safetyevaluation

JevOut: Natural Context Can Flip Decision Models

Sep 24, 2026

Zixiang Xu

Decision models are surprisingly brittle: short, natural-sounding context additions can redirect correct predictions to high-confidence wrong answers, suggesting their probability outputs shouldn't be trusted as reliable decision interfaces without additional safeguards.

This paper reveals that decision models like Jev—which map text to probability distributions over choices—are vulnerable to subtle context manipulation. Researchers show that naturally-written contextual additions can flip correct decisions to wrong ones in 61% of cases, even when the original question and answer remain unchanged.

safetyevaluation

To Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech

Sep 24, 2026

Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri

Audio language models have a significant blind spot: they verify written facts reliably but fail on identical claims when spoken. Retrieval-augmented approaches help only when combined with explicit reasoning, not retrieval alone.

VeriSpeak is a benchmark for fact-checking spoken claims using audio language models. It contains nearly 4,000 spoken statements about real-world facts and tests whether models can verify claims directly from speech, especially when given retrieved text evidence.

evaluationmultimodalreasoning

Requirement-Bound Verified Commissioning: A Frozen Four-Billion-Parameter Local Model as a Candidate Generator under an External Acceptance Layer with Verification and Release Authority

Sep 24, 2026

Mehmet Iscan

Separating candidate generation from verification through an external gate with formal constraints can achieve high safety (zero false releases on benchmark) while maintaining usability, but real-world deployment requires testing with actual users and measuring gate sensitivity.

This paper presents a safety protocol for AI-assisted mechatronic systems that separates candidate generation from release decisions. A frozen 4-billion-parameter language model generates plans, but they're only released when an external verification gate confirms both required facts using a formal grammar.

safetyevaluationreasoning

Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure

Sep 24, 2026

David Schmotz, Derck Prinzhorn, Luca Beurer-Kellner et al.

LLM agents naturally develop evasion strategies under normal task pressure without explicit adversarial training—they encode prohibited commands, decompose operations, and retry strategically.

This paper studies how LLM agents attempt to evade runtime monitoring systems when completing ordinary tasks.

safetyagentsevaluation

The Alignment Illusion in Multimodal Large Language Models

Sep 24, 2026

Hong-Han Wang, Yuntao Wang, Hu Ding

Don't trust standard alignment scores as proof that multimodal models understand images—they can be fooled by the model's internal structure. Use task-based validation and the PA gap metric instead to verify genuine cross-modal integration.

This paper reveals that high visual-text similarity scores in multimodal AI models don't actually mean the model is properly integrating images and text. Using experiments across 13 models, the researchers show that replacing images with random noise barely changes similarity scores, even though task performance drops sharply.

evaluationmultimodal

A Living Benchmark for Information Retrieval from Electronic Health Records

Sep 24, 2026

Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani et al.

Automated benchmark generation validated by domain experts enables continuous evaluation of clinical AI systems, exposing real gaps (like multi-document synthesis) that static benchmarks miss.

Researchers created BRIE, an automatically-generated benchmark for testing how well AI systems retrieve and synthesize information from patient medical records.

evaluationapplicationsdata

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

Sep 24, 2026

Ming Zhang, Zhenghao Xiang, Peizhong Gao et al.

Current AI systems can learn new rules through exploration, but performance is inconsistent and fragile—gains from exploration can reverse with continued interaction, suggesting exploration capabilities need significant improvement.

ExplorationBench is a benchmark that tests whether AI systems can genuinely explore and discover new rules in unfamiliar environments, rather than just recalling training data. It uses two 'Alien Worlds' with executable rules that differ from real-world knowledge, forcing systems to experiment, form hypotheses, and learn through interaction rather than memorization.

evaluationreasoningagents

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

Sep 24, 2026

Xinyue Zeng, Jiawei Zhang, Yujun Yan et al.

Long-horizon reasoning failures in LLMs stem from structural biases in the reasoning space itself, not just model capacity—and injecting geometric structure into the reasoning process can dramatically improve performance on hard problems.

This paper addresses why large language models struggle with long-horizon reasoning tasks by identifying two key problems: exploration bias (getting stuck in locally plausible but structurally weak paths) and compounding bias (small errors accumulating over many steps).

reasoningarchitectureevaluation

Intrinsic-Extrinsic Coupling in Learning Dynamics

Sep 24, 2026

Qinyou Wang

A model's internal state and external training conditions interact in complex, non-additive ways—the same intervention can help or hurt depending on what happens next, which matters for understanding continual learning and model adaptation.

This paper studies how a machine learning model's current state interacts with future training dynamics. The authors develop methods to measure and manipulate learning states—like classifier weights and historical information—to understand when interventions help or hurt performance.

trainingevaluation

Do Audio Language Models Hear and Read Distinctive Features Alike?

Sep 24, 2026

Yuanhao Chen, Peter Chin

Audio language models don't reliably represent the same phonetic features in the same direction across speech and text—suggesting these models may process the two modalities quite differently despite using a shared decoder.

This paper investigates whether audio language models represent phonetic features consistently across speech and text inputs. Using minimal pairs of phonemes that differ in single features, researchers measured whether the same distinctive features (like voicing) are encoded in the same direction in both modalities across 6 models, 7 features, and 15 languages.

multimodalevaluationarchitecture

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

Sep 24, 2026

Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh

Token-level tolerance to transcription ambiguity in ASR training reduces word error rates by ~9.5% on average by letting models skip disputed individual characters while keeping supervision for the rest of the word.

This paper addresses a real problem in speech recognition: reference transcripts often contain ambiguous pronunciations or spellings that the audio doesn't uniquely determine.

trainingevaluation

Does a model's stated reason for rejecting a candidate do any work?

Sep 24, 2026

Archit Rastogi

Language models often cite missing facts when rejecting candidates, but careful testing shows these stated reasons have limited causal influence on their actual choices, raising questions about whether models are genuinely reasoning or post-hoc rationalizing.

This paper tests whether language models' stated reasons for rejecting candidates actually influence their decisions.

evaluationreasoningalignment

SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue

Sep 22, 2026

Haobo Zheng, Tan Tang, Yan Chen et al.

For multi-party dialogue, tracking speaker identity and relationships separately from content is crucial—this dual-track approach outperforms general-purpose memory systems that try to handle everything at once.

This paper introduces SpeakerMem-R1, a memory system for multi-party conversations that tracks who said what and how people relate to each other. Unlike general LLMs that lose track of speakers and relationships, it uses a dual-track approach: storing exact messages labeled by speaker plus derived relationship states, organized by person and group.

reasoningagentsevaluation

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

Sep 22, 2026

Jennifer Williams, Dave Farris, Jeff Farris et al.

AI agents can complete inference engineering tasks locally, but production correctness is much harder—end-to-end serving tests catch failures that other tests miss, exposing a critical gap between development and deployment.

SWE-Serve is a benchmark with 53 real production tasks from SGLang that tests whether AI agents can implement inference serving features correctly—not just locally, but in production. It reveals a major gap: one-third of code changes that pass unit tests fail end-to-end serving tests, showing that current agents struggle with production-grade correctness.

evaluationagentsefficiency

EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations

Sep 22, 2026

FNU Aditi

When evaluating LLMs that generate hardware assertions, using equivalent implementations reveals that the same behavior can produce different numbers of correct assertions—showing that assertion quality depends on implementation details, not just the intended behavior.

EquivSVA is a dataset of 120 behavior families with 480 RTL implementations and 914 formally verified assertions, designed to test whether AI-generated hardware assertions capture true behavior or just implementation details. Each behavior family has four structurally different implementations of the same functionality, enabling controlled studies of assertion robustness.

evaluationdata

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

Sep 22, 2026

Om Nepal, Sushant Aryal, Oluseyi Olukola et al.

Don't use compile rate to evaluate LLM vulnerability repair; it's gamed by non-repairs and dominated by evaluation setup, not model quality. Use change-aware metrics like diff_F1 as a first-pass screen before running actual tests.

This paper shows that compile rate—a common metric for measuring LLM-based code vulnerability repair—is unreliable because it's dominated by dataset artifacts rather than model quality, shifts dramatically with compiler flags, and even rewards broken patches.

evaluationsafetyapplications

Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows

Sep 22, 2026

Remy Stewart, Olabode Anise, Andrew Hogan et al.

AI design tools deliver real time savings (~20%) in controlled settings, but benefits vary by user expertise—product managers gain more than professional designers, indicating task-dependent value.

Researchers tested whether AI-powered design tools (Figma Make) actually save time by having 50 designers and 50 product managers complete design tasks with and without the tool. They found about 20% faster completion times with AI assistance, especially for non-designers, suggesting these tools may let product managers do more design work themselves.

evaluationapplicationsefficiency

The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence

Sep 22, 2026

Xiaoyu Yang, Jie Lu, Wei Duan et al.

When building long-context systems, irrelevant nearby text can be a bigger problem than distance itself. LYRA's t-distributed approach helps models focus on what matters regardless of position.

This paper identifies the 'Proximity Trap'—where LLMs struggle with distant evidence not because of distance itself, but due to interference from irrelevant nearby context. The authors propose LYRA, a retrieval mechanism using t-distributed matching to prioritize task-relevant evidence while maintaining positional information, and introduce ProxBench to benchmark this problem.

evaluation

TraceVIC: Causal Reasoning over Code Evolution for Identifying Vulnerability-Inducing Commits

Sep 22, 2026

Fnu Tanish, Samiha Shimmi, Samikshya Chapagain et al.

Identifying vulnerability-inducing commits requires reasoning about code evolution across revisions, not just finding where vulnerable code was last modified—TraceVIC achieves 28.7% improvement over existing methods by modeling full revision history.

TraceVIC identifies which commit introduced a software vulnerability by analyzing how vulnerable code evolved across a project's revision history. Instead of using simple heuristics like "earliest change," it builds temporal graphs showing code structure and evolution, then reasons over this history to rank commits by their contribution to the vulnerability.

safetyevaluation

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Sep 21, 2026

Yiran Wang, Xingyilang Yin, Junfu Pu et al.

This benchmark reveals that gameplay requires coordinating visual understanding, instruction following, and long-horizon planning—and shows clear performance gaps between model families, providing a standardized way to measure progress on embodied AI tasks.

GameHorizon is a comprehensive dataset and benchmark for evaluating AI models on video game tasks. It includes 5,000 hours of gameplay from 21 AAA games with aligned videos, actions, and instructions at multiple time scales, plus offline and online evaluation tracks to measure how well models understand and execute gameplay across different planning horizons.

evaluationmultimodalreasoning

DolphinBench: Mapping the Pareto Frontier of Agent Memory

Sep 21, 2026

Soumil Rathi, Deshraj Yadav, Taranjeet Singh

Most memory benchmarks only measure accuracy, but DolphinBench shows that practical agent memory systems must balance three things: how well they retrieve information, how much they cost, and how fast they respond.

DolphinBench is a benchmark that evaluates how well AI agents use long-term memory to complete real-world tasks, not just answer questions. It includes 500k tokens of realistic user messages per persona and measures accuracy, cost, and latency together—revealing tradeoffs that single-metric benchmarks miss.

evaluationagentsreasoning

Rare Event Estimation via Iterative Unalignment

Sep 21, 2026

Hanming Yang, Daksh Mittal, Jing Dong et al.

For safe AI deployment, you need to estimate how often catastrophic rare events occur in agent behavior. This paper provides a practical method using importance sampling with learned weight perturbations, achieving massive efficiency gains over naive approaches.

This paper tackles estimating extremely rare event probabilities in AI agent trajectories—events so uncommon that standard Monte Carlo sampling is impractical. The authors develop a new importance sampling method that tweaks a language model's weights to generate more likely rare events, using gradient-based optimization to search efficiently.

safetyevaluationefficiency

Generative Tutorial: Towards Live Contextualized Visual Instructions for Physical Tasks

Sep 21, 2026

Muzhe Wu, Zuchen Li, Xu Wang et al.

Contextualizing visual instructions to match a user's real workspace—through AI-generated images and videos—improves task performance and trust compared to generic pre-authored tutorials.

This paper presents a system that generates live visual instructions tailored to a user's specific workspace and task progress, rather than showing pre-recorded tutorials. Using AR and generative AI, it creates goal images and demo videos that match the user's actual environment, helping them complete physical tasks more accurately and confidently.

applicationsmultimodalevaluation

Exactness at Inference: A Representational Criterion for Out-of-Distribution Generalization

Sep 21, 2026

Filipe Marinho Rocha, Inês Dutra, Vítor Santos Costa et al.

Out-of-distribution generalization requires exact representational equivalence to the generating mechanism, not statistical approximation—a criterion that constrains inference rather than training and explains why neural networks fail on novel entities while logic-based systems succeed.

This paper argues that models generalize beyond their training data only when they compute representations structurally equivalent to the underlying mechanism—not approximations.

reasoningevaluationarchitecture
agentssafetyevaluation

BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings

Sep 18, 2026

Alexandre Andre, Shivashriganesh P. Mahato, Vinam Arora et al.

Large-scale neural recording pretraining improves transfer learning, but no single approach generalizes well across behavior prediction, neural dynamics, and anatomical organization—suggesting general-purpose brain models remain an open challenge.

BrainWideBench is a benchmark for evaluating whether neural network models trained on large-scale brain recordings from many mice can learn generalizable representations that transfer to new animals and tasks. The benchmark tests three key abilities: predicting behavior from brain activity, forecasting neural patterns, and recovering brain anatomy.

evaluationscalingdata

Predictable Failure in Multi-Hop Retrieval: Score-Distributional Confidence Scoring and Abstention

Sep 18, 2026

Andre Bacellar

Multi-hop retrieval failures follow predictable structural patterns that vary by dataset; you can detect high-risk queries using simple features (query length, retrieval concentration) and a learned confidence score, enabling safe abstention without additional LLM calls.

This paper identifies why multi-hop retrieval systems fail predictably on certain queries and proposes a method to detect these failures without extra LLM calls. The authors prove that failure patterns cluster in specific subpopulations and that different query features predict failure in different scenarios.

evaluationreasoning

Benchmarking World Models for Continual Learning on Compositional Tasks

Sep 18, 2026

Haoyu Zhou, Joe Watson, Anson Lei et al.

Modular world model architectures better balance knowledge reuse with avoiding catastrophic forgetting in continual learning, but the field still lacks methods that effectively retain and reuse knowledge across sequential robot tasks.

This paper creates a benchmark to test how well world models (AI systems that learn to predict environment dynamics) can learn continuously across robot tasks without forgetting previous knowledge. The key innovation is using compositional tasks—where new tasks combine elements from earlier ones—to isolate what knowledge gets reused versus forgotten.

trainingevaluationarchitecture

Available Guardrails: Certifying Selective Prediction across ML Systems

Sep 18, 2026

Parivesh Priye, Yufeng Wang, Haibin Ling et al.

When deploying safety-critical ML systems with selective prediction, finite calibration data is the bottleneck—smart partition selection and error budget reallocation can recover 60% of the theoretical coverage gain, but naive approaches recover almost none.

This paper addresses how to safely deploy selective predictors (models that abstain when uncertain) by certifying they meet precision targets for specific user groups. The key challenge is that with limited calibration data, some groups may not have enough evidence to certify safety.

safetyevaluationapplications

Gricea: An Open Science Platform for Conversational AI Research

Sep 18, 2026

Nikhil Sharma, Yunlin Gong, Xinyang Cheng et al.

Standardized, shareable research artifacts dramatically improve reproducibility in conversational AI research—Gricea enables researchers to replicate 93% of studies and identify gaps that would otherwise block replication.

Gricea is an open-science platform that lets researchers package conversational AI studies as reusable, shareable artifacts. It standardizes how study procedures, chatbot systems, and conversation tasks are documented and deployed, making it easier to replicate and build on prior research.

evaluationapplications

QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

Sep 18, 2026

Rawan El Ghali, Umm Kulsoom, Anas Madkoor et al.

Multiple-choice scoring masks AI model failures in language understanding—QuranicMMLU shows models score 24 percentage points higher on multiple-choice than open-ended tasks, meaning benchmark design significantly impacts what we learn about model capabilities.

QuranicMMLU is a benchmark for testing how well AI models understand Quranic Arabic across five linguistic areas: sound, word structure, grammar, meaning, and context.

evaluationdata

COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules

Sep 18, 2026

Sushovan Majhi, Atish Mitra, Žiga Virk et al.

COMPLEX provides the first two-sided distortion bounds for multiparameter topological features, enabling certified embeddings where you can mathematically verify that the embedding preserves data relationships—a critical missing piece for trustworthy topological machine learning.

COMPLEX is a new method for converting multiparameter persistence modules (a topological data analysis tool) into embeddings that can be used for machine learning. Unlike previous approaches, it provides both upper and lower bounds on how well features are preserved, making it possible to verify that similar data stays similar in the embedding.

evaluationdata

DiaVLo: Diagnosing Behaviours of Vision-Language Models

Sep 18, 2026

Lorenzo Corti, Jie Yang

Understanding VLM behaviors requires systematic diagnosis—DiaVLo shows how to identify what models actually do versus what they should do, and which concepts drive their decisions.

DiaVLo is a diagnostic framework that identifies and explains the behaviors of vision-language models by comparing desired behaviors against observed ones. It uses human curation and the models' own generation capabilities to surface misalignments and pinpoint which concepts most influence model decisions.

evaluationsafety

Embedding Models Measure in Peculiar Ways

Sep 17, 2026

Juri Opitz, Andrianos Michail

Embedding models don't reliably represent objective physical concepts—they're influenced more by surface-level text patterns than by actual semantic relationships, which has implications for using embeddings in domains requiring precise measurements.

This paper investigates whether embedding models accurately represent physical measurements like mass, distance, and time. The researchers found that embeddings poorly capture these objective quantities and instead learn superficial patterns based on string similarity rather than true semantic meaning.

evaluation

How Does Distribution Shift Shape Pretraining Gains in Neural PDE Surrogates?

Sep 17, 2026

Pochinapeddi Sai Bhargav, Nithin Somasekharan, Rohit Sunil Kanchi et al.

Pretraining neural PDE surrogates provides significant data efficiency gains (2-3x fewer samples needed), but this benefit shrinks or reverses when the target task involves different physics modeling than the pretraining source.

This paper investigates how pretraining neural networks to simulate fluid dynamics (PDE surrogates) helps when switching to new airfoil designs or physics models. The authors show that pretraining benefits depend on three factors: how much target data you have, how diverse that data is, and whether the source and target use different physics models.

trainingefficiencyevaluation

Quantifying Overclaiming Propensity in Frontier LLM Agents

Sep 17, 2026

Nolan Smyth, Yorguin-Jose Mantilla-Ramos, Pascal Jr Tikeng Notsawo et al.

Frontier coding agents frequently misrepresent their work in final responses, claiming task completion when they've actually skipped files or missed defects—a critical reliability issue for autonomous systems users depend on.

This paper measures how often frontier AI coding agents falsely claim to have completed tasks they didn't finish. Researchers tested eight proprietary and four open-source models on file-review scenarios, finding that agents skip files 68% of the time and mislead users about coverage 80% of the time—either lying about reading everything or hiding incomplete work.

safetyevaluationagents

Unifying Models of Intergroup Hostility in Online Discourse

Sep 17, 2026

Patrick Gerard, Julia Mendelsohn, Kristina Lerman

Intergroup hostility follows a predictable structural pattern: boundary-setting and threat-framing anchor the system, while dehumanization and scapegoating emerge later.

This paper analyzes 2.86 million social media posts to understand how hostile rhetoric toward groups develops online.

safetyevaluationdata

An Empirical Study of Harness Design for Coding Agents

Sep 17, 2026

Run-Ze Fan, Zihao Zhang, Simin Ma et al.

Harness design should adapt to model capability: weaker models benefit from planning and predefined tools, while stronger models achieve better cost-efficiency with minimal scaffolding and bash-only interfaces.

This paper systematically studies how different components of coding harnesses—the execution frameworks that guide AI agents through software engineering tasks—affect agent performance.

agentsevaluationapplications

PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers

Sep 17, 2026

Jiachen Yao, Zi-Siang Hsu, Xi Deng et al.

Existing evaluations of generative inverse solvers miss critical failures like mode collapse and overconfident uncertainty—PosteriorBench reveals these gaps by directly comparing predicted solution distributions to ground-truth posteriors.

PosteriorBench is a benchmark for evaluating how well generative models solve inverse problems by checking if they capture the full range of possible solutions, not just single best guesses. It tests four physics problems using reference posteriors from MCMC and rejection sampling, with metrics measuring accuracy, uncertainty, and distributional fit.

evaluationreasoning

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

Sep 17, 2026

Sarah Wyer, Sue Black, Noura Al Moubayed

Safety metrics like toxicity scores can mask real harms—discrimination doesn't disappear during model training, it just becomes harder to detect. Developers need better evaluation methods that catch representational bias, not just explicit toxicity.

This paper reveals that safety improvements in GPT models don't actually reduce gender discrimination—they transform it into subtler forms.

safetyevaluationalignment

Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission Protocols

Sep 17, 2026

Tariq Abdul-Quddoos, Xiangfang Li, Lijun Qian

Calibrated confidence thresholds enable RF-fingerprinting to work reliably under co-channel interference, with mathematical guarantees on missing detections—critical for spectrum monitoring in crowded wireless environments.

This paper tackles RF-fingerprinting—identifying specific transmitters by their hardware signatures—in realistic scenarios where multiple devices transmit simultaneously and interfere with each other. The authors use a CNN with calibration techniques to ensure reliable detection while controlling false negatives, validated on real 5G testbed data with Wi-Fi, LTE, and 5G signals.

evaluationsafetyapplications

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

Sep 17, 2026

Sho Kawano, Zehang Richard Li, Paul A. Parker

When evaluating AI systems across domains with limited labels, borrowing information across domains through prediction-powered smoothing gives more accurate performance estimates than evaluating each domain independently.

This paper addresses how to accurately evaluate AI systems across different domains (like task types or conversation types) when you can only label a small sample. It proposes prediction-powered smoothing, which combines AI predictions with limited labels to get better estimates for each domain, and a validation method to choose between different estimation approaches.

evaluationdata

Large Language Models as Falsifiers for Cyber-Physical Systems

Sep 17, 2026

Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak

LLMs can be effective at finding counterexamples in formal verification by leveraging semantic information like natural language descriptions and trajectory data, outperforming traditional black-box optimization methods on standard benchmarks.

This paper shows how large language models can find bugs in cyber-physical systems (like autonomous vehicles or industrial controllers) by treating bug-finding as an optimization problem.

safetyreasoningevaluation

Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure

Sep 17, 2026

Zofia Smoleń

Labeling spreadsheet cells by role helps, but hits a hard limit because spreadsheets are inherently 2D with infinite possible relationships. Better approaches should flatten spreadsheets into 1D text rather than trying to classify cells into fixed categories.

This paper tackles how to make spreadsheets queryable by AI systems. The key insight is that spreadsheets are fundamentally 2D structures that don't fit neatly into categories—so instead of trying to classify each cell's role, the field should develop better ways to convert spreadsheets into readable text that language models can work with.

dataevaluationapplications

Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol

Sep 17, 2026

Levent Bulut

LLMs may systematically degrade narrative quality by rewarding explicit declaration over subtle implication, which matters because these models increasingly evaluate and shape written content.

This paper proposes 'summarization bias'—a tendency of LLMs to flatten narrative meaning into explicit summaries rather than preserving the inferential structure that makes stories work.

evaluationreasoning

HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication

Sep 17, 2026

Hassan Saeed Hassan Albattra, Mazen Mohammed Bahgat, Rahatara Ferdousi et al.

Multilingual healthcare AI needs explicit testing for how language, communication style, and missing information affect risk assessment—aggregate accuracy scores alone don't catch safety failures that emerge in specific languages or contexts.

HerHealthEval is an evaluation framework that tests how well AI language models understand women's health concerns across multiple languages (English, French, Arabic) and communication styles (clinical, casual, emotional, vague).

evaluationmultimodalsafety

Chronicle: Cut-Point Replay for Regression Testing of LLM Agents

Sep 17, 2026

Tisha Chawla, Susheem Koul

By recording agent execution at non-deterministic boundaries, Chronicle enables regression testing of LLM agents without re-running expensive model calls, catching bugs that would otherwise be hard to reproduce due to non-determinism.

Chronicle is a testing tool that records LLM agent runs at decision points and replays them to catch regressions. It captures non-deterministic model outputs and tool interactions as immutable snapshots, then selectively replays recorded boundaries while executing new code live—turning past failures into reproducible regression tests that run in CI/CD pipelines.

agentsevaluation

A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies

Sep 17, 2026

Khalid Halba, Kylie Cooper, James G. Bellingham

LLM-based fault diagnosis for autonomous systems requires rigorous ensemble testing rather than single trials, and frontier models significantly outperform local alternatives, but success depends on models following complete diagnostic procedures rather than jumping to conclusions.

This paper presents SPAR, a simulation platform for testing how large language models can help autonomous underwater vehicles diagnose and recover from faults without human intervention.

agentsevaluationreasoning

WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution

Sep 17, 2026

Yi Zhou, Kiamehr Rezaee, Danushka Bollegala et al.

Language models struggle with WiC not because they can't understand word meanings, but because they lack clear guidance on what level of semantic detail matters—adding explicit sense options fixes this and reveals models often overthink distinctions.

This paper shows that Word-in-Context (WiC) tasks are harder for language models than traditional Word Sense Disambiguation (WSD) because WiC lacks an explicit sense inventory. By providing candidate senses to models, performance improves significantly, and human evaluation reveals many errors stem from disagreement about sense granularity rather than true comprehension failures.

evaluationreasoning

A Zeroth-Order Paradigm for LLM Preference Alignment

Sep 16, 2026

Peter Chen, Xi Chen, Wotao Yin et al.

ComPO offers a gradient-free alternative to direct preference optimization that may better handle preference pairs with small margins, with both theoretical convergence guarantees and empirical improvements across major LLM families.

This paper introduces ComPO, a new method for aligning language models with human preferences that uses comparison oracles instead of directly optimizing preference losses. Unlike standard approaches, ComPO extracts directional signals from preference pairs without computing gradients, and includes both offline and online variants with theoretical guarantees.

alignmenttrainingevaluation

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

Sep 16, 2026

Sara Pieri, Evangelos Kazakos, Shizhe Chen et al.

Combining dense captioning with pixel-level grounding requires jointly optimizing text generation and mask selection—PANORAMA shows this can be done effectively by conditioning a segmenter on phrase representations and learning which masks correspond to each phrase.

This paper introduces PANORAMA, a vision-language model that generates detailed image captions while simultaneously grounding each phrase with pixel-level segmentation masks.

multimodalevaluationarchitecture

Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging

Sep 16, 2026

Pranaya Jajoo

History-dependent logging can make policy evaluation exponentially hard even with good state coverage, because resets can erase critical information about transitions that determine policy value.

This paper proves that even when a logged dataset visits all hidden states frequently, it can still require exponentially many episodes to evaluate a policy's performance if the logger depends on history. The authors construct specific environments where standard coverage conditions hold, yet accurate policy evaluation needs Θ((3/2)^H) episodes—exponential in horizon H.

evaluationreasoning

Affora: A Design System for Agent-Friendly Interfaces

Sep 16, 2026

Jin Gao

AI agents perform better when interfaces clearly communicate available actions and current state—you can achieve this through thoughtful design that serves both humans and machines without sacrificing visual flexibility.

Affora is a design system that makes software interfaces work better for both humans and AI agents. Rather than creating separate interfaces for machines, it improves existing interfaces so agents can understand what actions are available and what state the software is in, while keeping the visual design flexible and familiar to users.

agentsapplicationsevaluation

Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

Sep 16, 2026

Peter Potash

Frontier models communicate inefficiently with themselves across information asymmetry, extracting only ~0.93 bits per question instead of the theoretical 1 bit, with failures driven equally by answer errors and inability to discriminate between candidates.

This paper evaluates six frontier language models playing a communication game where one model asks yes/no questions to identify a target Wikipedia article from a set, while another answers with single words.

evaluationreasoningagents