ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2567 papers9 this month12 topics
AllTraining 46Reasoning 42Efficiency 33Evaluation 32Agents 27Architecture 18Multimodal 17Applications 16Data 9Safety 9scaling 5Alignment 2

Oct 5 – Oct 11(2)

Base Models Can Reason By Taking a Cue From Training Data

Oct 5, 2026

Sophie L. Wang, Amil Dravid, Rulin Shao et al.

Base models already contain reasoning capabilities encoded in their training data—you can unlock them by conditioning on the right token cues, without needing expensive RL fine-tuning.

This paper shows that base language models can achieve reasoning performance comparable to RL-trained models by using specific starting tokens (like "Okay" or "Alright") that trigger learned associations from training data. The authors demonstrate they can create new reasoning cues through data interventions and trace these effects back to specific document types in the training set.

trainingreasoningdata

PlotGround: Grounding Plot Digitization in Real Scientific Figures and Their Source Data

Oct 5, 2026

Yaohui Zhang, Binxu Li, Haoyi Duan et al.

Current multimodal models can approximate values from scientific figures but struggle with precision; providing source data instead of figures dramatically improves accuracy (90% to 97.4%) while reducing computational cost, suggesting a practical path for scientific data extraction.

PlotGround is a benchmark for evaluating how well AI models can extract numerical values from scientific figures.

Sep 28 – Oct 4(7)

RNADyn: A Benchmark for Generating and Understanding RNA Dynamics

Oct 2, 2026

Yiming Huang, Lennart Bastian, Hanqun Cao et al.

A unified deep learning approach can both generate realistic RNA dynamics trajectories and predict dynamics fingerprints from static structures, bridging two previously separate tasks and improving physical accuracy through explicit physical constraints.

This paper introduces RNADynBench, a large-scale benchmark of 2,585 RNA molecular dynamics simulations, and RNADynNet, a unified model that generates realistic RNA trajectories and extracts dynamics information from single structures.

dataarchitectureevaluation

LESSER: Post-Training Data Selection with Output-Layer Gradients

Oct 2, 2026

Lyuxin David Zhang, Eric Wong, Surbhi Goel et al.

You can select effective training data for LLMs using only output-layer gradients from forward passes, cutting computation costs by ~10× compared to full-gradient methods without sacrificing downstream performance.

This paper proposes LESSER, a method for selecting high-quality training data for large language models by using only output-layer gradients instead of full-parameter gradients. By leveraging cheaper forward passes rather than expensive backward passes, LESSER reduces computational cost by 9.7× for supervised fine-tuning while maintaining performance comparable to full-gradient selection methods.

Sep 21 – Sep 27(6)

MexHat: A Dataset for Hate Speech Detection in Mexican Spanish Videos

Sep 25, 2026

Itzel Tlelo-Coyotecatl, Hugo Jair Escalante

Most hate speech detection datasets focus on English; this dataset enables models to learn culturally-specific patterns in Mexican Spanish video content, which is essential since hate speech is deeply tied to local context and language nuances.

MexHat is a new video dataset with ~1,000 annotated clips for detecting hate speech in Mexican Spanish. It addresses the lack of non-English resources by capturing linguistic and cultural context specific to Mexico, with annotations for both general categories (offensive vs. hate speech) and fine-grained hate speech subtypes.

datamultimodalsafety

A Living Benchmark for Information Retrieval from Electronic Health Records

Sep 24, 2026

Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani et al.

Automated benchmark generation validated by domain experts enables continuous evaluation of clinical AI systems, exposing real gaps (like multi-document synthesis) that static benchmarks miss.

Researchers created BRIE, an automatically-generated benchmark for testing how well AI systems retrieve and synthesize information from patient medical records.

Sep 14 – Sep 20(9)

BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings

Sep 18, 2026

Alexandre Andre, Shivashriganesh P. Mahato, Vinam Arora et al.

Large-scale neural recording pretraining improves transfer learning, but no single approach generalizes well across behavior prediction, neural dynamics, and anatomical organization—suggesting general-purpose brain models remain an open challenge.

BrainWideBench is a benchmark for evaluating whether neural network models trained on large-scale brain recordings from many mice can learn generalizable representations that transfer to new animals and tasks. The benchmark tests three key abilities: predicting behavior from brain activity, forecasting neural patterns, and recovering brain anatomy.

evaluationscalingdata

Particle Competition and Cooperation for Robust Graph Convolutional Network Learning Under Label Noise

Sep 18, 2026

Fabricio Breve

Pre-processing noisy labels with a lightweight particle-based algorithm before GCN training significantly improves robustness to label corruption while being faster than other robust methods.

This paper proposes PCC+GCN, a method that cleans noisy labels in graph data before training a Graph Convolutional Network. It uses a particle-based algorithm to identify and fix mislabeled nodes, then trains GCN on the refined labels. The approach is faster and more accurate than existing robust GCN methods across multiple datasets with different types of label noise.

Sep 7 – Sep 13(6)

Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact

Sep 10, 2026

Masahiro Kato, Daiki Honma, Taka Kato

Businesses can now quantify the ROI of appearing in generative AI outputs by combining visibility data with causal inference, enabling marketing decisions in an AI-driven world where traditional metrics don't apply.

This paper introduces a framework to measure how generative AI systems (like ChatGPT) affect business outcomes when companies optimize their visibility in AI-generated answers.

applicationsevaluationdata

CausalArena: Benchmarking Causal Discovery in the Foundation Model Era

Sep 10, 2026

Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang et al.

No single causal discovery method generalizes across different benchmark types; evaluation results depend heavily on which test scenario you use, making diverse benchmarks essential for fair comparison.

CausalArena is a comprehensive benchmark for evaluating causal discovery methods—tools that learn cause-and-effect relationships from data. It addresses a critical problem: existing benchmarks use different test scenarios, making it hard to compare methods fairly.

Aug 31 – Sep 6(12)

A Deep Generative Model for Synthesizing Labeled Wireless Signals

Sep 4, 2026

Yuxiao Li, Keke Hu, Santiago Mazuelas et al.

Synthetic wireless signals generated by GANs can effectively replace expensive real-world data collection for training wireless sensing models, reducing labeling costs while maintaining physical realism.

This paper introduces IIns-GAN, a deep learning method that generates realistic labeled wireless signals for training machine learning models.

datatrainingapplications

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

Sep 4, 2026

Dain Kim, Eungi Cho, Kyumin Kim et al.

Open-source models can match much larger models at multi-step API calling by training on execution-verified data—showing that smart data synthesis matters more than model size for tool-use tasks.

This paper addresses a real-world problem: open-source AI models struggle when they need to call multiple government APIs in sequence to complete tasks. The authors create KOPA-Bench, a benchmark of 145 real Korean government API tasks, and introduce EDGE, a method that learns which API outputs can feed into other APIs by actually testing them live.

Aug 24 – Aug 30(5)

SWE-Prime: Fewer Trajectories, Better Performance

Aug 27, 2026

Dewu Zheng, Ruizhe Ye, Yanlin Wang et al.

Filtering training data for code-fixing tasks at both trajectory and step levels beats training on all successful examples—quality and relevance of supervision matter more than quantity.

SWE-Prime improves how AI models learn to fix software bugs by being smarter about which training examples to use. Instead of training on all successful bug fixes, it filters trajectories (problem-solving paths) at two levels—first selecting high-quality complete solutions, then identifying which individual steps within those solutions are actually worth learning from.

trainingdataapplications

Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation

Aug 27, 2026

Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong et al.

By modeling reactions as electron rearrangements instead of molecular graph edits, MAELLE achieves competitive accuracy while naturally explaining reaction mechanisms and maintaining robustness on out-of-distribution chemistry.

This paper presents MAELLE, a machine learning model that predicts chemical reactions by tracking how electrons move between atoms, rather than just predicting final products.

Aug 17 – Aug 23(14)

Across-Design Uncertainty in Short Pricing Panels: Evidence from Simulated Price Trajectories

Aug 21, 2026

Pedro Cadahia Delgado

When estimating prices from sparse movement data, most uncertainty comes from which price trajectory could have occurred, not from fitting a single trajectory—and standard statistical methods don't capture this.

This paper analyzes estimation uncertainty in short pricing datasets where few distinct price movements occur despite many observations. Using simulated data, the author shows that most estimation error (97.6%) comes from variation across different possible price trajectories rather than uncertainty within a single trajectory.

evaluationdata

Inducing Task Models from Computer-Use Traces

Aug 20, 2026

Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen et al.

By converting raw computer activity into structured task models with goal hierarchies and control flow, TMI enables AI agents to learn realistic work procedures and organizations to audit and reuse task knowledge from employee activity traces.

This paper presents Task Model Induction (TMI), a method that automatically discovers and structures how people actually work on computers by analyzing screenshots and input logs.

agents

Aug 10 – Aug 16(14)

Split the Labor: Separating Evidence Interpretation from Decision Aggregation

Aug 14, 2026

Zhelun Wu

When combining evidence from multiple sources, separate the task of interpreting each source from aggregating those interpretations—use structured tuples and calibrated scoring rather than simple concatenation and vote counting.

This paper separates evidence interpretation from decision aggregation in multi-source reasoning systems. Instead of concatenating sources into one prompt, the authors propose a structured evidence tuple (hypothesis, reliability, rationale, provenance) and show how to properly combine interpretations using calibrated log-likelihood ratios.

reasoningevaluationdata

RecipeNet: A Hierarchical Transformer for Recipe Data

Aug 14, 2026

Pin-Yen Huang, Sachin Chhabra, Prasanth Sai Gouripeddi et al.

Hierarchical structure matters: representing recipes as nested sequences of structured steps, rather than flattened tables, lets models learn procedural dependencies and field interactions that improve performance on real-world synthesis and manufacturing tasks.

RecipeNet is a hierarchical Transformer model designed to learn from recipe data—ordered sequences of steps with structured fields—used in materials science, pharmaceuticals, and manufacturing. Unlike traditional tabular methods that flatten this data, RecipeNet captures both field interactions within steps and dependencies across steps, achieving better performance on recipe-based tasks.

Aug 3 – Aug 9(10)

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

Aug 6, 2026

Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi et al.

Automated feature engineering for clinical data is feasible when grounded in clinical guidelines and evidence trails, but requires careful validation and auditing to ensure reliability in real-world healthcare settings.

Researchers built an automated system (nMAS) to extract and engineer features from fragmented heart-failure patient records in electronic health records. The system combines multi-agent AI with clinical guidelines to generate interpretable features, reducing manual work that typically consumes 39-45% of data scientists' time.

applicationsdataagents

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Aug 6, 2026

Fanzhe Meng, Guoxin Chen, Jiale Zhao et al.

Training agents on tasks calibrated to be appropriately difficult (not too easy, not impossible) using multiple solver feedback produces better generalization than manually authored or single-solver validated tasks.

CalibForge automatically creates training tasks for AI agents by using multiple solvers to identify tasks that are challenging but solvable—the 'learnable zone.' It revises candidate tasks based on solver disagreement and performance patterns, then trains agents on these calibrated tasks, achieving significant improvements on code and repository understanding benchmarks.

Jul 27 – Aug 2(15)

Differentially Private Nonparametric Modal Learning with Applications to Regression and Clustering

Jul 31, 2026

Arkajyoti Bhattacharjee, Arnab Auddy

You can privately estimate where data clusters (density modes) with theoretical guarantees on both privacy and accuracy, achieving near-optimal statistical rates that balance the privacy-utility tradeoff.

This paper develops methods for finding density modes (peaks in probability distributions) while guaranteeing differential privacy—a mathematical constraint that limits what can be learned about individual data points. The authors propose DP-GRAMS, which uses noisy gradient ascent on a privately estimated score function, and prove it recovers all modes with near-optimal accuracy.

safetyevaluationdata

Freeze, Then Select: Structured Field Adapters and Stability-Validated Weak Selection for PDE Discovery from Sparse Observations

Jul 31, 2026

Juncheng Zhong, Chenghuang Shen, Jianfeng Liu et al.

Decoupling field reconstruction from equation selection—by freezing a learned field representation and then selecting terms via stability-validated weak-form analysis—improves PDE discovery from sparse observations compared to end-to-end neural approaches.

This paper tackles PDE discovery from sparse data by separating field reconstruction from equation selection. The authors develop a freeze-then-select method that first trains a neural adapter to reconstruct the continuous field, then uses stability-validated weak selection to identify the correct differential terms.

evaluationmultimodaldata
trainingefficiencydata

Higher-Order Molecular Grammars for Generative and Foundation Models in Chemistry

Oct 1, 2026

Yiming Huang, Yujie Zeng, Vijay Prakash Dwivedi et al.

HGR enables sequence models to generate chemically valid molecules with perfect validity while capturing complex molecular topology, achieving top performance on generation and property prediction tasks without the computational cost of explicit higher-order encodings.

This paper introduces Higher-order Grammar Representation (HGR), a new way to represent molecules that captures complex structural features like ring systems by converting them into sequences of grammar rules. Unlike existing methods that struggle with computational overhead, HGR makes these structures compatible with standard sequence models while guaranteeing valid molecules.

architecturedataapplications

Finetuning with Sampling: SFT Learns Better Than You Think

Oct 1, 2026

Aayush Karan, Sitan Chen, Yilun Du

SFT isn't inherently worse than RL for posttraining—the gap comes from data distribution mismatch. By reshaping data to be more on-policy before training, SFT can generalize better and forget less than strong RL baselines.

This paper shows that supervised finetuning (SFT) can match or beat reinforcement learning for posttraining if you transform the training data first. The authors use an MCMC sampling algorithm to gradually shift off-policy expert demonstrations toward on-policy trajectories that a reference model can actually learn from.

trainingdata

Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows

Oct 1, 2026

Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing et al.

Current frontier AI models struggle with real enterprise data work: the best model scores 95+ on only 35% of tasks, revealing a major gap between text-to-SQL benchmarks and actual data agent capabilities needed for production systems.

Argo-Bench is an evaluation framework with 210 realistic data science tasks that test AI agents on enterprise-scale workflows.

evaluationagentsdata

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Oct 1, 2026

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc et al.

Spatially grounded self-distillation with synthetic data can teach multimodal models better visual reasoning that transfers to real-world tasks, without requiring human annotations or external teachers.

This paper improves multimodal AI models by having them learn from a smarter version of themselves that receives spatial hints about where to look in images. Using synthetic scenes with automatic object labels, the approach trains models to understand spatial relationships without human annotation, and surprisingly, these improvements transfer to real-world vision tasks.

trainingmultimodaldata

Scalable, Transferable Meta-network for Data Selection Requires a Different Loss (and Why the Obvious Choice is Problematic)

Oct 1, 2026

Zilin Du, Bowen Yang, Boyang Albert Li

When scaling data selection with neural networks, the standard loss function causes poor generalization—a new loss function (PVM) that matches predicted values pointwise solves this and transfers better across datasets and model scales.

This paper addresses data selection for training large language models by proposing TESS, a framework that uses a neural network to score and select training examples. Unlike existing meta-learning approaches that assign per-sample weights, TESS uses a novel loss function (Pointwise Value Matching) that avoids optimization instability and improves generalization to new datasets and model sizes.

trainingdataefficiency
evaluationapplicationsdata

ARGUS: Role-Aware Event Knowledge Graphs for U.S. Employment-Discrimination Complaints

Sep 24, 2026

Sriram Kannan, Swetha Saseendran, Vishnu Vardhan Reddy Kandi et al.

Structured event graphs from legal documents improve reasoning tasks like classification and QA, but only when you already have the right documents—they don't help with initial retrieval.

ARGUS builds structured Event Knowledge Graphs from employment-discrimination legal complaints using a 5W1H schema and LLMs. The system extracts facts, organizes events with temporal and causal relationships, and merges them into document-level graphs. Testing shows these graphs improve claim classification and help answer legal questions when relevant documents are already retrieved.

dataapplicationsreasoning

FleXray: Universal Clinical X-ray Segmentation

Sep 22, 2026

Victor Ion Butoi, Vivek Gopalakrishnan, John V. Guttag et al.

You can train powerful medical imaging models without expensive manual annotation by simulating realistic training data from existing 3D datasets—FleXray segments full-body X-rays and works on real clinical data despite being trained entirely on synthetic images.

FleXray is a generalist AI model that segments 60 anatomical structures in clinical X-rays across the entire body. Rather than manually labeling thousands of X-rays, the researchers built a physics-based simulator that generates realistic synthetic X-rays from existing 3D CT scans, then trained the model on these simulations.

trainingdataapplications

EquivSVA: A Formally Verified Dataset of Behavioral Assertions Across Equivalent RTL Implementations

Sep 22, 2026

FNU Aditi

When evaluating LLMs that generate hardware assertions, using equivalent implementations reveals that the same behavior can produce different numbers of correct assertions—showing that assertion quality depends on implementation details, not just the intended behavior.

EquivSVA is a dataset of 120 behavior families with 480 RTL implementations and 914 formally verified assertions, designed to test whether AI-generated hardware assertions capture true behavior or just implementation details. Each behavior family has four structurally different implementations of the same functionality, enabling controlled studies of assertion robustness.

evaluationdata

onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

Sep 21, 2026

Lei Yang, Mengyin Liu, Jia Wang et al.

Token-level correction during annotation is significantly faster than full rewriting and produces on-policy training data that preserves the model's natural generation patterns while providing precise supervision signals.

onPanda is an interactive annotation tool that helps create training data for AI models by letting annotators correct responses token-by-token. Instead of rewriting entire outputs, annotators find the first mistake, fix it, and let the model regenerate from that point.

trainingdataalignment
trainingdata

QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

Sep 18, 2026

Rawan El Ghali, Umm Kulsoom, Anas Madkoor et al.

Multiple-choice scoring masks AI model failures in language understanding—QuranicMMLU shows models score 24 percentage points higher on multiple-choice than open-ended tasks, meaning benchmark design significantly impacts what we learn about model capabilities.

QuranicMMLU is a benchmark for testing how well AI models understand Quranic Arabic across five linguistic areas: sound, word structure, grammar, meaning, and context.

evaluationdata

COMPLEX: A Closed-Form Certified Embedding of Multiparameter Persistence Modules

Sep 18, 2026

Sushovan Majhi, Atish Mitra, Žiga Virk et al.

COMPLEX provides the first two-sided distortion bounds for multiparameter topological features, enabling certified embeddings where you can mathematically verify that the embedding preserves data relationships—a critical missing piece for trustworthy topological machine learning.

COMPLEX is a new method for converting multiparameter persistence modules (a topological data analysis tool) into embeddings that can be used for machine learning. Unlike previous approaches, it provides both upper and lower bounds on how well features are preserved, making it possible to verify that similar data stays similar in the embedding.

evaluationdata

Unifying Models of Intergroup Hostility in Online Discourse

Sep 17, 2026

Patrick Gerard, Julia Mendelsohn, Kristina Lerman

Intergroup hostility follows a predictable structural pattern: boundary-setting and threat-framing anchor the system, while dehumanization and scapegoating emerge later.

This paper analyzes 2.86 million social media posts to understand how hostile rhetoric toward groups develops online.

safetyevaluationdata

Prediction-Powered Smoothing and Validation for Disaggregated AI Evaluation

Sep 17, 2026

Sho Kawano, Zehang Richard Li, Paul A. Parker

When evaluating AI systems across domains with limited labels, borrowing information across domains through prediction-powered smoothing gives more accurate performance estimates than evaluating each domain independently.

This paper addresses how to accurately evaluate AI systems across different domains (like task types or conversation types) when you can only label a small sample. It proposes prediction-powered smoothing, which combines AI predictions with limited labels to get better estimates for each domain, and a validation method to choose between different estimation approaches.

evaluationdata

Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure

Sep 17, 2026

Zofia Smoleń

Labeling spreadsheet cells by role helps, but hits a hard limit because spreadsheets are inherently 2D with infinite possible relationships. Better approaches should flatten spreadsheets into 1D text rather than trying to classify cells into fixed categories.

This paper tackles how to make spreadsheets queryable by AI systems. The key insight is that spreadsheets are fundamentally 2D structures that don't fit neatly into categories—so instead of trying to classify each cell's role, the field should develop better ways to convert spreadsheets into readable text that language models can work with.

dataevaluationapplications

PAA: The Probabilistic Allen Algebra: A Generative and Complete Probabilistic Extension of Allen's Interval Relations

Sep 17, 2026

Julian Eggert

Probabilistic Allen algebra lets you reason about uncertain temporal relations from language and data by deriving relation probabilities from distributions over interval boundaries, rather than assigning scores arbitrarily—making temporal reasoning more grounded in uncertainty.

This paper extends Allen's interval algebra—a system for reasoning about temporal relationships—to handle uncertainty. Instead of crisp yes/no answers about whether one event is "before" another, it models time points and interval boundaries as probability distributions (Gaussians), allowing expressions like "roughly during" or "just before" to have graded meanings.

reasoningdata

Dreaming the Sound of Contact: Leveraging Video and Audio Generation for Zero-Shot Force-Aware Manipulation and Data Generation

Sep 16, 2026

Guanhua Ji, Tianyu Li, Dayoon Suh et al.

Combining generated video with generated audio allows robots to infer not just motion but also the forces needed for contact-heavy tasks—something video alone cannot provide.

This paper shows how robots can learn contact-rich manipulation tasks by generating both video and audio together. The system uses the loudness of generated contact sounds to create force profiles that guide the robot's movements, enabling tasks like pushing and grasping that require precise force control. The approach works without task-specific training data.

agentsmultimodaldata
evaluationreasoningdata

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Sep 10, 2026

Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi

Arabic speech AI was severely underrepresented in multilingual models—this project provides the dataset, trained models, and evaluation tools needed to build general-purpose Arabic speech systems from scratch.

Nuha-Speech addresses the lack of Arabic speech AI models by creating a comprehensive framework: a 1.5M-sample Arabic speech question-answering dataset, fine-tuned models based on Qwen-Omni, and evaluation benchmarks. This work establishes foundational infrastructure for building Arabic-capable speech AI systems despite limited available resources.

multimodaldataevaluation

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing

Sep 10, 2026

Pruthwik Mishra, Rudra Trivedi, Avi Patel et al.

Contextual transformer models like MuRIL and XLM-RoBERTa can effectively identify which language each token belongs to in code-mixed text, outperforming traditional monolingual approaches.

This paper tackles language identification in code-mixed text (where people mix multiple languages in one message) by treating it as a sequence labeling task. Researchers fine-tuned transformer models designed for Indian languages on three language pairs and released benchmark datasets with manually annotated examples to help future work on this social media problem.

datatraining

Dynamic language model representations for multi-objective reaction optimisation

Sep 10, 2026

Joshua W. Sin, David Ming Segura, Bojana Ranković et al.

Language models can learn task-specific representations of chemical reactions directly from text, enabling more efficient multi-objective optimization than hand-crafted chemical descriptors—achieving high-yield synthesis conditions with under 3% of the design space explored.

Researchers developed a method to optimize chemical reactions across multiple goals (yield, selectivity, safety) by using a fine-tuned language model to learn dynamic representations of reaction components from text descriptions.

applicationsdata

Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech

Sep 10, 2026

Chibuzor Okocha, Christan Earl Grant

Standard speech recognition metrics mask critical failures on code-switched speech—you need switch-aware diagnostics to understand where systems actually break down, especially for low-resource languages like Yoruba.

This paper evaluates how well modern speech recognition systems handle code-switched speech (mixing English and Yoruba). Standard metrics like word error rate hide the real problems: systems struggle much more with Yoruba than English, and fail especially at language switches.

evaluationdata
agentsdatatraining

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

Sep 4, 2026

Ankit Goyal, Jaideep Ray

Memory portability isn't automatic: compressed formats couple tightly to specific models and fail asymmetrically during upgrades, while structured knowledge graphs and raw histories are more robust. Always test migrations in both directions and keep original source data for recovery.

When you upgrade an AI agent's model, its memory often breaks in unexpected ways. This paper tests four memory storage formats—raw text, chunked retrieval, compressed notes, and structured knowledge graphs—to see which survive model swaps. Fixed schemas stay reliable, but compressed notes and retrieval systems lose accuracy unpredictably depending on which direction you upgrade.

agentsevaluationdata

LexFlip: A Dissociation Diagnostic for Legal Meaning Preservation Metrics

Sep 4, 2026

Gaurab Baral

Standard semantic similarity metrics are unreliable for legal text simplification because they can't distinguish between preserving words and preserving legal meaning—a problem that requires new evaluation approaches beyond token overlap.

This paper exposes a critical flaw in how we measure whether simplified legal text preserves meaning. Current metrics fail because they conflate surface-level word overlap with actual legal meaning.

evaluationdata

Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views

Sep 3, 2026

Joseph Lee, Yidi Huang, Dokyoon Kim et al.

LLMs learn knowledge more effectively when exposed to multiple reformulations of the same concept rather than just repeated text, suggesting that data diversity—not just quantity—is fundamental to pre-training success.

This paper investigates how LLMs learn knowledge during pre-training by studying the role of 'auxiliary views'—different reformulations of the same knowledge. Through controlled experiments, the researchers show that diverse representations of knowledge improve learning better than simple repetition, even when the total number of tokens stays the same.

trainingdata

Last Translation Benchmark

Sep 3, 2026

Vilém Zouhar, Niyati Bafna, Mukund Choudhary et al.

Standard translation benchmarks are saturating and automatic metrics are unreliable—this benchmark uses human-authored hard cases with explicit failure rules to provide reproducible evaluation that actually identifies what models get wrong.

This paper introduces the Last Translation Benchmark, a curated collection of challenging translation examples (text, images, audio, video) designed to expose weaknesses in state-of-the-art machine translation models.

evaluationdata

Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments

Sep 3, 2026

Jie Wu, Zhenru Zhang, Beichen Zhang et al.

Instead of building environments from scratch, you can extract them from existing agent trajectories, then use those environments to generate diverse training tasks at scale—turning single frozen demonstrations into many verifiable, interactive learning scenarios.

Terminal-Universe reconstructs executable coding environments from agent trajectories by replaying file operations and synthesizing missing dependencies. It then generates new tasks from these environments—both variations of original tasks and cross-workspace queries—and extends them into multi-turn interactions with simulated user feedback.

trainingagentsdata

AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application

Sep 2, 2026

Wenxin Jiang, Xuyang Wang, Yuxiao Wu

AI can reliably measure hard-to-survey characteristics at the individual level, but only when you have rich background data available—it's most useful for filling gaps in existing datasets rather than replacing surveys entirely.

Researchers developed AICOME, a framework for validating whether AI-generated measures of occupational and social characteristics can accurately recover both individual and group-level effects in statistical models.

evaluationdata

Closing Cost-Quality Gap in Document VLMs: Difficulty-Aware Data Curation and Quality-Adjusted Deployment Economics

Sep 1, 2026

Maksim Evdokimov, Matvey Ivanov, Dmitrii Tsiupin et al.

Smart data curation and cost-aware deployment economics can make small VLMs competitive with much larger models for document processing—the key is selecting training examples strategically and measuring real-world costs including verification and correction.

This paper presents a practical document understanding system that extracts structured data from documents at a fraction of the cost of human annotation or larger models.

efficiencydataapplications

Configurable Semantic Chunking for Biomedical Information Extraction in Retrieval-Augmented Generation

Aug 31, 2026

Riya Ahuja, Tim Kacprowski, Roya Shiasi Sardoabi

Semantic chunking that respects entity and relationship boundaries significantly outperforms fixed-size chunking for biomedical RAG, especially when relation cues are explicit in the text.

This paper improves biomedical information extraction in RAG systems by replacing fixed-size text chunking with a configurable semantic chunking framework. The approach preserves important entities and relationships by using trigger-centered chunking and hierarchical relation resolution, achieving 8.4 F1 points improvement on biomedical relation extraction benchmarks.

dataapplications

OntoAligner-Ensemble: Voting-Based Fusion across Heterogeneous Ontology Alignment Techniques

Aug 31, 2026

Hamed Babaei Giglou, Sören Auer, Peio Popov et al.

Combining different types of ontology alignment methods through voting-based ensemble learning reliably improves results, with the best approach depending on whether you prioritize precision (mix different paradigms) or overall F1-score (use multiple LLMs).

This paper presents OntoAligner-Ensemble, a framework that combines predictions from different ontology alignment methods (string-based, knowledge graph embeddings, and LLM-based) using voting and selection strategies. Testing across biomedical and other domains shows that mixing diverse alignment approaches consistently improves precision-recall balance and often beats individual methods.

evaluationdata

LLM Post-Training as Brownfield Maintenance: An Industrial Perspective on Dataware Engineering

Aug 31, 2026

Gopi Krishnan Rajbahadur, Amir M. Ebrahimi, Boyuan Chen et al.

Post-training in production is about engineering discipline around data mixtures and yield metrics, not algorithmic breakthroughs—small improvements in converting training data into usable supervision can compound into significant model gains.

This paper treats LLM post-training as industrial maintenance work, not research. Teams inherit a deployed model and must improve it within strict compute budgets without breaking existing capabilities.

trainingdataapplications
reasoningapplicationsdata

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

Aug 27, 2026

Maayan Sharon, Tom Hope

Retrieval systems need to understand different types of scientific inspiration—not just finding similar papers, but finding papers that generalize ideas, concretize them, or solve stated problems.

RATIO is a benchmark for retrieving scientific papers that can inspire research at different levels of abstraction. It defines three types of retrieval: finding approaches to solve problems, generalizing to broader concepts, or narrowing down to concrete implementations.

evaluationdataapplications

CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases

Aug 27, 2026

Sil Hamilton, Albert Yu Sun, Oscar J. Romero et al.

LLMs perform much worse on corporate Q&A tasks at realistic scale (thousands of documents) than on smaller benchmarks, and CorporateBench provides a standardized way to measure this performance gap without exposing real company data.

CorporateBench is a large-scale Q&A benchmark for testing how well language models handle real-world corporate document collections. It includes 230,000+ human-validated documents organized into synthetic companies of varying sizes, with questions requiring both information extraction and knowledge base querying.

evaluationdataapplications

RCMN: Understanding Misleadingness in Influential Public Discourse

Aug 27, 2026

Peiling Yi

Misleadingness in public discourse is diverse and often subtle (exaggeration, omission, unsupported inference), and while AI can sometimes predict how readers will interpret misleading content, identifying *how* it misleads requires rich contextual and evidential grounding.

This paper introduces RCMN, a framework for understanding how public discourse misleads readers through framing, omission, and context rather than just false claims.

evaluationdatasafety
data
applications

MidTool: Mid-training Data Synthesis for Agentic Tool Use

Aug 20, 2026

Fengqing Jiang, Yite Wang, Boyi Liu et al.

Tool-use capabilities in language models improve significantly when trained during mid-training with targeted synthetic data, rather than waiting until post-training—similar to how math and reasoning skills benefit from dedicated training phases.

MidTool is a data synthesis pipeline that creates training data for teaching language models to use tools effectively during mid-training (the stage between pretraining and fine-tuning).

trainingagentsdata

Physical-Support Confidence Sets for Highly Coherent Dictionaries

Aug 20, 2026

Guan-Ju Peng

Highly coherent learned dictionaries can assign different physical meanings to the same sparse representation; you need to account for dictionary uncertainty and check which physical interpretations survive across all calibration-compatible alternatives.

When learning dictionaries from data to represent signals sparsely, the selected atoms may seem physically meaningful but could be arbitrary artifacts of the calibration data—especially when many different dictionaries fit equally well.

evaluationdata

Electronic Navigational Chart Change Classification

Aug 20, 2026

Jacob Arndt, Abhishek Potnis, Alexandre Sorokine

Machine learning can effectively automate safety-critical geospatial data review by encoding spatial context and attribute information, achieving 5-7% accuracy gains over baseline approaches and scaling beyond manual workflows.

This paper automates the classification of changes to Electronic Navigational Charts (ENCs)—geospatial datasets critical for maritime safety—by converting complex vector data into structured formats and using gradient-boosted trees. The method achieves 90-94% accuracy on real operational data, reducing manual review burden and improving consistency in identifying safety-critical chart updates.

dataevaluationapplications

Multi-Method Causal Evidence Synthesis: Ranking Candidate Drivers by Convergent Cross-Method Evidence from Observational Data

Aug 20, 2026

Manish Gupta, Dipanjan De

When inferring causality from observational data, combining evidence across multiple methods—even those with different assumptions—identifies likely causal relationships more reliably than trusting any single method alone.

This paper presents MCES, a framework that combines outputs from eleven causal inference methods across different mathematical traditions to rank which variables most likely cause outcomes in observational data.

evaluationdata

Decoding silent reading from non-invasive EEG

Aug 20, 2026

Ingo Marquardt, Anthilia Alchanat, Priyanka Jain

Silent reading produces decodable word-level information in EEG that improves with more training data—this opens a scalable path to studying how the brain processes language without relying on unreliable self-reports of inner speech.

Researchers decoded which words people were silently reading from brain activity (EEG), using 49 hours of data from one participant. They trained a neural network to match EEG signals with word embeddings from a language model, achieving above-chance accuracy on 240,000 word presentations.

evaluationmultimodaldata

OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models

Aug 20, 2026

Nikita Khudov

Domain-specific benchmarks built from verified sources reveal that LLM performance gaps aren't just about model size—reasoning capabilities, self-preference bias, and access to reference materials dramatically change accuracy, with some models gaining 33+ percentage points when given context.

OenoBench is a wine-domain benchmark with 3,266 multiple-choice questions built from verified facts extracted from government registries and peer-reviewed sources.

evaluationdata

Learned, Then Lost: A Measured Single-Example Counterfactual in Pre-training

Aug 19, 2026

Zachary Speck, Asa Shepard

Individual training examples can be learned and then completely forgotten during pre-training, leaving virtually no measurable impact on final model behavior or internal structure, suggesting that what matters for final performance is the aggregate signal, not individual data points.

Researchers trained 32 GPT-2 models from scratch and injected a single training example at peak learning rate to measure its actual impact. The example was learned immediately but completely forgotten by the end of training, leaving no detectable trace in the final model's weights, geometry, or performance—despite moving the model within its loss basin during training.

trainingdata

Comment-level Topic Drift Analysis in the Reddit Corpus

Aug 19, 2026

Steven Morse, Daniel Runfola, Trenton W. Ford

Political and social topics show measurable semantic drift in online discourse over time, detectable through embedding-space analysis—a technique that could help track how language and meaning evolve around contentious issues.

This paper analyzes how topics shift and evolve in Reddit discussions over 16 years by tracking semantic embeddings of 12.7 billion comments. Using pretrained language models and unsupervised clustering, the researchers show that politically charged topics drift significantly in meaning over time, while stable domains like sports remain consistent.

dataevaluationapplications

An Analytical-Prior Framework for Data-Efficient Prediction of Sound-Reduction Frequencies in Rectangular Side-Branch Helmholtz Resonators

Aug 17, 2026

Jiaming Li

When you have a fast but imperfect analytical model and limited expensive simulation data, teaching a neural network to correct the analytical model's errors—or pre-training it on the analytical model first—can cut your data requirements dramatically.

This paper shows how to make machine learning models more data-efficient by combining cheap analytical equations with expensive high-fidelity simulations. Using Helmholtz resonators as a test case, the authors demonstrate two approaches: learning to correct analytical predictions, or distilling analytical knowledge into a neural network before fine-tuning with limited simulation data.

dataefficiencytraining

zLend: A Dual-Scope Cash-Flow Reconstruction Framework for On-Chain Credit Underwriting

Aug 17, 2026

Girish G N, Ashutosh Sahoo, Akshay SP et al.

On-chain lending needs to separate a wallet's total token holdings from its liquid, spendable balance—two wallets with identical net worth can have very different repayment capacity depending on which assets they actually hold.

zLend is a framework that assesses borrower creditworthiness in decentralized lending by analyzing on-chain transaction history. It reconstructs daily wallet balances from token transfers in two ways—using only stablecoins and using all tokens—to distinguish between total wealth and spendable cash.

applicationsdataevaluation

Time-Aware Validation of Machine Learning Fuel Consumption Models: Evidence from 1\,Hz Operational Data, CCGS \textit{Sir Wilfrid Laurier}

Aug 17, 2026

Samarasimha Reddy Chittamuru, Ayhan Akinturk, Allison Kennedy et al.

When validating ML models on time-series data like operational sensor readings, use time-aware cross-validation instead of random splits—random splits create unrealistic performance estimates that won't hold in real deployment.

This paper reveals a critical flaw in how machine learning models for ship fuel consumption are validated: most studies use random train-test splits that leak temporal information and give overly optimistic results.

evaluationdata

GEO-Flag: Detecting and Measuring GEO-Optimized Web Content

Aug 17, 2026

Junjie Chu, Ye Leng, Mingjie Li et al.

Generative search engines are vulnerable to content optimization tactics that inflate the visibility of low-authority or false information; systematic detection methods are now possible but require careful design to avoid relying on author-based shortcuts.

This paper introduces GEO-Flag, a system for detecting web pages optimized for generative search engines (like Google's AI Overviews).

safetyevaluationdata
architecturedata

Generating Benchmark Health Data Using a Tabular Diffusion Transformer

Aug 14, 2026

Hao Yan, Lisa Pilgram, Dan Liu et al.

You can now generate realistic synthetic tabular health data from multiple heterogeneous sources by standardizing them statistically first, then using diffusion transformers to learn and reproduce their patterns.

This paper presents a method for generating synthetic health data from multiple different database tables with varying structures. It works in two stages: first converting diverse tables into a standardized statistical format, then using a diffusion transformer to learn patterns and generate new synthetic tables.

dataarchitectureevaluation

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Aug 13, 2026

Fanfei Li, Jana Zeller, Manuel Prada-Corral et al.

Training on a carefully curated, grade-level-appropriate curriculum creates a sandbox for studying knowledge acquisition with clear boundaries—useful for understanding how models learn and what happens when you try to teach them new concepts.

Researchers created LittleLeaner, a 5B-parameter language model trained on an 88B-token curriculum limited to U.S. Grade 5 material, to study how models acquire knowledge under controlled conditions. Unlike models trained on messy web data, LittleLeaner has clear, interpretable knowledge boundaries, making it easier to understand what the model knows and how it learns new information.

trainingdataevaluation

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

Aug 13, 2026

Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina et al.

You can build competitive frontier-level language models at 1B parameters using only openly licensed data, making it feasible for researchers and organizations to develop ethical AI without relying on scraped or restricted datasets.

Mimir v1 is a 1-billion-parameter language model trained entirely on permissible (legally and ethically sourced) data that achieves competitive performance with much larger models. It uses a Hierarchical Reasoning Model architecture and excels at English, math, code, and Danish tasks—showing that high-quality open-source models don't require massive proprietary datasets.

trainingdataefficiency

Measuring Task-Agnostic Training Data Influence Across Language Model Pretraining

Aug 13, 2026

Yuto Nishida, Hirokazu Kiyomaru, Yusuke Oda et al.

You can measure training data importance during pretraining by tracking parameter trajectories, revealing that different data types matter at different training stages—literature early, STEM later—without needing task-specific validation sets.

This paper introduces a new method to measure how much training data influences language model development without needing to pick specific downstream tasks. Instead of testing on particular benchmarks, the researchers measure influence by tracking how each piece of training data pushes the model toward its final parameters.

trainingdataevaluation

TabSOM: A tabular-to-image encoding method based on self-organizing maps

Aug 13, 2026

David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara et al.

Self-organizing maps can encode tabular data more effectively than simpler dimensionality reduction by capturing feature relationships alongside values, improving both predictive performance and model interpretability.

TabSOM converts tabular data into images using self-organizing maps, preserving both feature values and relationships between features. Unlike existing methods that only encode individual feature values, TabSOM captures feature interactions as spatial patterns, enabling vision models to achieve better performance while remaining interpretable.

dataarchitectureevaluation

Sparse Orthogonal Regression Technique: A Spectral Framework for Equation Discovery, Approximation, and Integration

Aug 13, 2026

Sabin Roman, Ljupco Todorovski, Saso Dzeroski

SORT shifts equation discovery from brittle library selection to basis design: by learning sparse coefficients in well-chosen orthogonal bases, it provides a more stable intermediate representation that gracefully degrades under noise and sampling sparsity.

SORT is a machine learning technique that learns compact mathematical representations of dynamical systems from noisy, irregularly sampled data by fitting sparse coefficients in orthogonal basis expansions.

reasoningdataarchitecture

Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

Aug 13, 2026

Serli Kopar, Sam Gijsen, Abner Hernandez et al.

Speech-based disease detection models may not be learning genuine disease characteristics but rather dataset-specific patterns, raising serious concerns about their reliability for real-world clinical use.

This paper investigates whether speech models trained to detect Parkinson's disease actually learn disease-specific patterns or just exploit dataset quirks.

evaluationmultimodaldata

Wasserstein Filtering: A Sample Selection Method for Robust Distribution Learning

Aug 13, 2026

Yikai Xu, Zhao Chen, Jian Huang

You can robustly learn from contaminated data by selecting the subset of samples that maximizes Wasserstein distance from the full dataset—this works as a model-agnostic preprocessing tool before training any model.

This paper introduces Wasserstein Filtering, a method to clean contaminated datasets by selecting samples whose distribution is most different from the full dataset. The approach uses optimal transport theory to identify and remove outliers, with theoretical guarantees and practical algorithms that work as a preprocessing step for any downstream task.

dataevaluation

Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

Aug 12, 2026

Avijit Roy, Proma Roy

AI systems for underrepresented languages fail at the infrastructure level—in data collection, tokenization, and deployment design—long before model training begins. Fixing this requires treating offline-first design and linguistic diversity as core infrastructure priorities, not afterthoughts.

This paper reveals how AI infrastructure systematically disadvantages speakers of underrepresented languages like Bengali before models are even trained.

dataapplications

A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement

Aug 12, 2026

Bryan Torres, Daniel Riofrío, José Vega-Sánchez et al.

Lightweight, domain-specific NLP models can effectively detect fraud signals in real-world government data without expensive infrastructure—useful for building practical oversight tools in resource-constrained settings.

This paper develops a hybrid NLP pipeline to detect accusatory language in public procurement comments from Ecuador's procurement system.

applicationsdata

ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls

Aug 11, 2026

Chen Lyu, Xingwei Tan, Simon Cullen et al.

This work demonstrates how to generate sensitive, realistic training data for abuse detection by modeling VAWG as temporally unfolding multi-turn conversations rather than isolated toxic sentences, enabling better downstream safety systems.

ConVAWG is a framework for generating synthetic multi-turn dialogues depicting Violence Against Women and Girls scenarios. Using retrieval-grounded methods, persona seeds, crime definitions, and real case reviews, it creates realistic abuse scenarios with controlled toxicity while respecting privacy constraints that prevent releasing real conversation data.

datasafetyapplications

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

Aug 11, 2026

Changhao Xiang, Shangyu Xing, Zhen Wu et al.

By interleaving visual objects directly into text during pretraining, you can teach multimodal models object-level grounding 12x more efficiently than traditional image-text pair training.

This paper introduces MultiModal Code-Switching (MMCS), a new way to train vision-language models by replacing words in text with their corresponding visual objects. Instead of just pairing whole images with descriptions, MMCS explicitly shows the model which objects match which words, making training much more efficient—achieving the same performance with 12x less data.

multimodaltrainingdata
trainingdataagents

Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

Aug 6, 2026

Donna Hooshmand, Shubham Shahi, Cameron Barrie et al.

Automating semantic schema construction lets non-technical users query databases without expert help, and TYTAN's hybrid symbolic-LLM approach achieves production-ready accuracy by knowing when to ask humans for clarification.

TYTAN automatically builds semantic schemas for relational databases by combining symbolic analysis with LLM inference to identify entities, relationships, and data roles. It asks targeted questions when ambiguous, achieving 100% coverage and correctness on real-world databases—eliminating the manual work that currently bottlenecks data analysis tools.

dataapplicationsreasoning

Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints

Aug 6, 2026

Omid Bazgir, Md Nasir, Jacob Hoffman et al.

Synthetic clinical benchmarks need explicit realism optimization separate from utility validation—passing utility checks alone doesn't guarantee a benchmark reflects real-world data patterns, which matters for training reliable healthcare AI agents.

This paper addresses a critical gap in synthetic clinical benchmarks: they can pass utility checks while remaining structurally unrealistic. The authors develop methods to improve benchmark realism (measured by data missingness patterns, actionability, and population alignment) while maintaining the utility thresholds required for downstream AI systems.

evaluationdatasafety

OTLesMix: Wasserstein Barycenter and Optimal Transport Map for Synthetic Lesion Generation with Diverse Shapes and Locations

Aug 6, 2026

Robin Trombetta, Carole Lartizien

Using optimal transport to blend medical images creates more realistic and varied synthetic lesions than traditional mixing strategies, leading to better segmentation model performance.

This paper presents OTLesMix, a data augmentation method that uses optimal transport and Wasserstein barycenters to generate synthetic medical images with diverse lesion shapes and locations. Tested on brain lesion segmentation, it improves model performance by 2.9-6.6 Dice points compared to baseline and outperforms existing mixing-based augmentation methods.

datatrainingevaluation

RxnCLF: Contrastive Transformation-Aware Reaction Foundation Model for Improved Reactivity Prediction

Aug 6, 2026

Yiting Zheng, Cheng Fang, Anthony Donofrio et al.

Using a unified graph representation of reactions instead of separate reactant/product graphs enables better learning of chemical transformations, leading to more accurate yield predictions and a foundation model that could generalize across diverse reaction types.

RxnCLF is a self-supervised learning framework that represents chemical reactions as unified graphs (condensed reaction graphs) to better capture how molecules transform.

trainingapplicationsdata

Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training

Aug 5, 2026

Damien Sileo, Valentin Lacombe, Dimitri Kachler

Procedural datasets for training reasoning models need careful design beyond correctness: difficulty calibration, compact targets, and rigorous auditing (combining model review, human judgment, and testing) significantly impact training utility.

This paper introduces Reasoning Core, a collection of 50 procedural generators that create verifiable reasoning problems across diverse domains (math, logic, planning, code, etc.). The authors compare their dataset with three alternatives using completion-supervised fine-tuning on 3B models, showing Reasoning Core achieves better performance on reasoning benchmarks.

trainingdatareasoning

OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling

Aug 5, 2026

Indraneil Paul, Falko Helm, Goran Glavaš et al.

Training language models on code contexts that span multiple related files—rather than isolated snippets—dramatically improves their ability to handle long contexts and understand complex codebases, even when this data makes up a small fraction of total training.

OctoLong is a pipeline that creates long, dependency-rich code contexts by automatically retrieving related code files across repositories.

trainingdata

Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation

Aug 4, 2026

Junhao Chen, Mingjin Chen, Jingjia Mao et al.

Music tokenization design is more important than model size for text-to-music generation—a small model with the right token representation outperforms massive models with poor representations, challenging the field's scaling assumptions.

This paper investigates how music tokenization—the way music is converted into discrete symbols for language models—affects text-to-music generation quality. By fixing model size, data, and training approach while swapping only the tokenization scheme, researchers found that representation choice matters far more than model scale.

architecturedataevaluation

Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation

Aug 4, 2026

Seyed Kahaki, Shijie Li, Weijie Chen et al.

When generating synthetic medical images, use domain-specific evaluation metrics rather than generic ones—and prioritize generating diverse data over pixel-perfect realism for downstream medical AI tasks.

This paper evaluates how well synthetic histopathology images generated by diffusion models work for medical AI tasks. The authors show that standard image quality metrics (FID, IS) designed for natural images fail for medical images, and propose using pathology-specific metrics instead.

evaluationdataapplications
reasoningdataevaluation

Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics

Jul 31, 2026

Yimin Chen, Brian Fricke, Bo Shen et al.

FDD-ON solves interoperability problems in HVAC fault diagnosis by creating a shared vocabulary and logical structure that lets different diagnostic tools, datasets, and applications understand each other.

This paper introduces FDD-ON, a structured knowledge framework (ontology) for understanding and diagnosing faults in HVAC systems. It standardizes how different systems describe equipment problems, symptoms, and impacts, enabling AI tools and building management systems to share diagnostic information reliably.

applicationsdata

AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

Jul 30, 2026

Bing Yan, Gregory Wolfe, Stefano Martiniani et al.

Instead of searching papers, search claims: AskChem lets you find specific findings with built-in source verification, reducing manual work for literature synthesis and improving AI agent reliability by 12% on citation accuracy.

AskChem is a search infrastructure that converts chemistry papers into atomic, provenance-tracked claims rather than returning full documents. Scientists and AI agents can search across 2.4M claims from 147K papers, with each claim linked to its source DOI and exact evidence location, enabling faster synthesis of findings across multiple papers.

applicationsdata

Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments

Jul 30, 2026

Haomin Qi, Xingliang Wang, Xuanqi Gao et al.

By mining repository history and reconstructing code states, you can generate verified, executable coding tasks at scale—reducing the manual effort of creating training data for code-generation agents while maintaining realistic development scenarios.

Change2Task automatically converts pull requests from repository history into executable coding tasks for training AI agents.

trainingdataagents

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

Jul 30, 2026

Manyi Wang, Junjielong Xu, Pinjia He

When using SWE-bench benchmarks to evaluate LLM coding abilities, be aware that over 1 in 10 test cases have misaligned problem statements and solutions—PAIChecker can automatically identify these problematic cases to improve benchmark reliability.

This paper identifies a critical quality issue in SWE-bench-like benchmarks used to evaluate AI coding abilities: 13.6% of PR-Issue pairs are misaligned, meaning the issue description doesn't actually match the code changes.

evaluationdata

Doubly Robust Functional Representation Learning for Longitudinal Causal Inference with Irregular Histories

Jul 30, 2026

Mengfei Ran, Yifeng Shen, Ruijie Guan

When analyzing longitudinal medical data with irregular measurements, you can use neural encoders to create summaries that maintain the statistical properties needed for valid causal inference—representation error behaves like ordinary estimation error under explicit conditions.

This paper addresses causal inference from messy medical data with irregular measurements (lab values, vital signs) taken at different times. The authors propose DR-FRL, a method that converts these fragmented histories into meaningful summaries while preserving statistical guarantees.

evaluationdata

ScaFE: Data-Efficient Scar Classification with LLM-Generated Clinical Feature Programs

Jul 30, 2026

Ruman Wang, Hangting Ye

LLMs can transfer medical knowledge into auditable, locally-executable feature programs instead of making direct predictions—improving data efficiency, privacy, and reproducibility in medical image classification.

ScaFE uses an LLM to generate executable programs that measure clinical scar features from photos, rather than asking the model to diagnose directly. These programs run locally, protecting patient data, and feed structured features into a simple Random Forest classifier.

applicationsefficiencydata

AI systems and the reproduction of (standard) language ideologies in World Englishes

Jul 30, 2026

Kingsley Ugwuanyi

AI systems aren't neutral—they embed and amplify existing power structures around which English dialects are considered 'legitimate,' with real consequences for speakers of non-standard varieties who are increasingly mistaken for AI.

This paper examines how AI language models reflect and reinforce biases toward 'standard' English while marginalizing non-dominant varieties spoken globally. Using examples from training data, model design, and public discourse, it shows how AI systems reproduce language ideologies that privilege English from wealthy countries while treating other Englishes as suspect or AI-like.

safetydataalignment

Beyond Sentiment: Structured Information Extraction from Financial News

Jul 30, 2026

Daohan Zhu, Sitong Ge, Ruofei Wang et al.

Financial sentiment analysis misses critical predictive signals; extracting structured semantic dimensions (event type, impact scope, temporal horizon, confidence) alongside sentiment improves stock prediction accuracy and reveals that sentiment-return relationships are highly nonlinear.

This paper shows that financial news contains multiple independent information dimensions beyond sentiment—like event type, impact scope, and time horizon—that together predict stock movements better than sentiment alone.

applicationsdata

Improving Mental Health Screening and Early Risk Detection in Spanish

Jul 30, 2026

Andreu Casamayor-Segarra, Vicent Ahuir, Antonio Molina-Marco et al.

Domain-specific Spanish models combined with ICE's automatic relabeling can detect mental health risks earlier and more reliably than general approaches, addressing a critical gap in non-English mental health AI tools.

This paper tackles mental health screening in Spanish by creating specialized language models and a new method called Incremental Context Expansion (ICE) that automatically identifies when enough social media messages accumulate to signal a mental health disorder. The approach reduces the time needed to detect problems while maintaining accuracy, with all models made publicly available.

applicationstrainingdata

Towards Autonomous Aircraft Surveillance from Nanosatellites through On-Board Inference and Generative Data Augmentation

Jul 30, 2026

Antonio Delgado-Rosa, David Muñoz-Valero, Enrique Adrian Villarrubia-Martin et al.

On-device inference on resource-constrained satellites combined with synthetic data generation can enable autonomous, real-time aircraft detection without overwhelming downlink capacity—shifting satellites from passive data collectors to active decision-makers.

This paper tackles satellite-based aircraft detection by running AI inference directly on small satellites (CubeSats) instead of sending raw images to Earth, and uses AI-generated synthetic images to train better detection models. The approach balances limited satellite bandwidth with scarce training data, enabling real-time autonomous surveillance from orbit.

efficiencyapplicationsdata

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

Jul 30, 2026

Jia Yu, Yan Zhu, Yili He et al.

Routine clinical reports contain rich supervision signals—you can extract finding-to-frame correspondence at scale to train medical vision-language models without expensive frame-level annotation, turning existing documentation into training data.

EndoCLIP is a vision-language model trained on 280,000 colonoscopy reports to link clinical findings with individual video frames.

multimodaldata

From Classification to Regression: Using a Fruitfly to Solve Equations

Jul 29, 2026

Shady E. Ahmed, Panos Stinis

You can solve regression problems by storing local patterns and using similarity matching instead of training large global models—this is faster, uses less memory, and works well for scientific data that repeats in certain regions.

This paper proposes a novel regression method inspired by how fruitflies sense their environment. Instead of building complex global models, the approach stores a library of representative local patterns and makes predictions by finding similar patterns to a query and combining their responses.

efficiencyreasoningdata

DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data

Jul 27, 2026

Zhen Huang, Yikun Wang, Shijie Xia et al.

Instead of applying uniform data processing rules, adapting the cleaning strategy per example—deciding what operation each piece of data needs—improves LLM pretraining efficiency and downstream performance.

DataOrchestra is a framework that customizes data processing for each example in pretraining, rather than applying one fixed strategy to all data. An orchestrator decides whether to drop, keep, or clean each data chunk, and if cleaning is needed, selects specific operations like editing or rewriting.

trainingdataefficiency