Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.
Sophie L. Wang, Amil Dravid, Rulin Shao et al.
Base models already contain reasoning capabilities encoded in their training data—you can unlock them by conditioning on the right token cues, without needing expensive RL fine-tuning.
This paper shows that base language models can achieve reasoning performance comparable to RL-trained models by using specific starting tokens (like "Okay" or "Alright") that trigger learned associations from training data. The authors demonstrate they can create new reasoning cues through data interventions and trace these effects back to specific document types in the training set.
Yaohui Zhang, Binxu Li, Haoyi Duan et al.
Current multimodal models can approximate values from scientific figures but struggle with precision; providing source data instead of figures dramatically improves accuracy (90% to 97.4%) while reducing computational cost, suggesting a practical path for scientific data extraction.
PlotGround is a benchmark for evaluating how well AI models can extract numerical values from scientific figures.
Yiming Huang, Lennart Bastian, Hanqun Cao et al.
A unified deep learning approach can both generate realistic RNA dynamics trajectories and predict dynamics fingerprints from static structures, bridging two previously separate tasks and improving physical accuracy through explicit physical constraints.
This paper introduces RNADynBench, a large-scale benchmark of 2,585 RNA molecular dynamics simulations, and RNADynNet, a unified model that generates realistic RNA trajectories and extracts dynamics information from single structures.
Lyuxin David Zhang, Eric Wong, Surbhi Goel et al.
You can select effective training data for LLMs using only output-layer gradients from forward passes, cutting computation costs by ~10× compared to full-gradient methods without sacrificing downstream performance.
This paper proposes LESSER, a method for selecting high-quality training data for large language models by using only output-layer gradients instead of full-parameter gradients. By leveraging cheaper forward passes rather than expensive backward passes, LESSER reduces computational cost by 9.7× for supervised fine-tuning while maintaining performance comparable to full-gradient selection methods.
Itzel Tlelo-Coyotecatl, Hugo Jair Escalante
Most hate speech detection datasets focus on English; this dataset enables models to learn culturally-specific patterns in Mexican Spanish video content, which is essential since hate speech is deeply tied to local context and language nuances.
MexHat is a new video dataset with ~1,000 annotated clips for detecting hate speech in Mexican Spanish. It addresses the lack of non-English resources by capturing linguistic and cultural context specific to Mexico, with annotations for both general categories (offensive vs. hate speech) and fine-grained hate speech subtypes.
Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani et al.
Automated benchmark generation validated by domain experts enables continuous evaluation of clinical AI systems, exposing real gaps (like multi-document synthesis) that static benchmarks miss.
Researchers created BRIE, an automatically-generated benchmark for testing how well AI systems retrieve and synthesize information from patient medical records.
Alexandre Andre, Shivashriganesh P. Mahato, Vinam Arora et al.
Large-scale neural recording pretraining improves transfer learning, but no single approach generalizes well across behavior prediction, neural dynamics, and anatomical organization—suggesting general-purpose brain models remain an open challenge.
BrainWideBench is a benchmark for evaluating whether neural network models trained on large-scale brain recordings from many mice can learn generalizable representations that transfer to new animals and tasks. The benchmark tests three key abilities: predicting behavior from brain activity, forecasting neural patterns, and recovering brain anatomy.
Fabricio Breve
Pre-processing noisy labels with a lightweight particle-based algorithm before GCN training significantly improves robustness to label corruption while being faster than other robust methods.
This paper proposes PCC+GCN, a method that cleans noisy labels in graph data before training a Graph Convolutional Network. It uses a particle-based algorithm to identify and fix mislabeled nodes, then trains GCN on the refined labels. The approach is faster and more accurate than existing robust GCN methods across multiple datasets with different types of label noise.
Masahiro Kato, Daiki Honma, Taka Kato
Businesses can now quantify the ROI of appearing in generative AI outputs by combining visibility data with causal inference, enabling marketing decisions in an AI-driven world where traditional metrics don't apply.
This paper introduces a framework to measure how generative AI systems (like ChatGPT) affect business outcomes when companies optimize their visibility in AI-generated answers.
Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang et al.
No single causal discovery method generalizes across different benchmark types; evaluation results depend heavily on which test scenario you use, making diverse benchmarks essential for fair comparison.
CausalArena is a comprehensive benchmark for evaluating causal discovery methods—tools that learn cause-and-effect relationships from data. It addresses a critical problem: existing benchmarks use different test scenarios, making it hard to compare methods fairly.
Yuxiao Li, Keke Hu, Santiago Mazuelas et al.
Synthetic wireless signals generated by GANs can effectively replace expensive real-world data collection for training wireless sensing models, reducing labeling costs while maintaining physical realism.
This paper introduces IIns-GAN, a deep learning method that generates realistic labeled wireless signals for training machine learning models.
Dain Kim, Eungi Cho, Kyumin Kim et al.
Open-source models can match much larger models at multi-step API calling by training on execution-verified data—showing that smart data synthesis matters more than model size for tool-use tasks.
This paper addresses a real-world problem: open-source AI models struggle when they need to call multiple government APIs in sequence to complete tasks. The authors create KOPA-Bench, a benchmark of 145 real Korean government API tasks, and introduce EDGE, a method that learns which API outputs can feed into other APIs by actually testing them live.
Dewu Zheng, Ruizhe Ye, Yanlin Wang et al.
Filtering training data for code-fixing tasks at both trajectory and step levels beats training on all successful examples—quality and relevance of supervision matter more than quantity.
SWE-Prime improves how AI models learn to fix software bugs by being smarter about which training examples to use. Instead of training on all successful bug fixes, it filters trajectories (problem-solving paths) at two levels—first selecting high-quality complete solutions, then identifying which individual steps within those solutions are actually worth learning from.
Nguyen Xuan-Vu, Octavian Susanu, Daniel Armstrong et al.
By modeling reactions as electron rearrangements instead of molecular graph edits, MAELLE achieves competitive accuracy while naturally explaining reaction mechanisms and maintaining robustness on out-of-distribution chemistry.
This paper presents MAELLE, a machine learning model that predicts chemical reactions by tracking how electrons move between atoms, rather than just predicting final products.
Pedro Cadahia Delgado
When estimating prices from sparse movement data, most uncertainty comes from which price trajectory could have occurred, not from fitting a single trajectory—and standard statistical methods don't capture this.
This paper analyzes estimation uncertainty in short pricing datasets where few distinct price movements occur despite many observations. Using simulated data, the author shows that most estimation error (97.6%) comes from variation across different possible price trajectories rather than uncertainty within a single trajectory.
Yucheng Jiang, Zora Zhiruo Wang, Ruishi Chen et al.
By converting raw computer activity into structured task models with goal hierarchies and control flow, TMI enables AI agents to learn realistic work procedures and organizations to audit and reuse task knowledge from employee activity traces.
This paper presents Task Model Induction (TMI), a method that automatically discovers and structures how people actually work on computers by analyzing screenshots and input logs.
Zhelun Wu
When combining evidence from multiple sources, separate the task of interpreting each source from aggregating those interpretations—use structured tuples and calibrated scoring rather than simple concatenation and vote counting.
This paper separates evidence interpretation from decision aggregation in multi-source reasoning systems. Instead of concatenating sources into one prompt, the authors propose a structured evidence tuple (hypothesis, reliability, rationale, provenance) and show how to properly combine interpretations using calibrated log-likelihood ratios.
Pin-Yen Huang, Sachin Chhabra, Prasanth Sai Gouripeddi et al.
Hierarchical structure matters: representing recipes as nested sequences of structured steps, rather than flattened tables, lets models learn procedural dependencies and field interactions that improve performance on real-world synthesis and manufacturing tasks.
RecipeNet is a hierarchical Transformer model designed to learn from recipe data—ordered sequences of steps with structured fields—used in materials science, pharmaceuticals, and manufacturing. Unlike traditional tabular methods that flatten this data, RecipeNet captures both field interactions within steps and dependencies across steps, achieving better performance on recipe-based tasks.
Soorya Ram Shimgekar, Michelle Hu, Dorisa Shehi et al.
Automated feature engineering for clinical data is feasible when grounded in clinical guidelines and evidence trails, but requires careful validation and auditing to ensure reliability in real-world healthcare settings.
Researchers built an automated system (nMAS) to extract and engineer features from fragmented heart-failure patient records in electronic health records. The system combines multi-agent AI with clinical guidelines to generate interpretable features, reducing manual work that typically consumes 39-45% of data scientists' time.
Fanzhe Meng, Guoxin Chen, Jiale Zhao et al.
Training agents on tasks calibrated to be appropriately difficult (not too easy, not impossible) using multiple solver feedback produces better generalization than manually authored or single-solver validated tasks.
CalibForge automatically creates training tasks for AI agents by using multiple solvers to identify tasks that are challenging but solvable—the 'learnable zone.' It revises candidate tasks based on solver disagreement and performance patterns, then trains agents on these calibrated tasks, achieving significant improvements on code and repository understanding benchmarks.
Arkajyoti Bhattacharjee, Arnab Auddy
You can privately estimate where data clusters (density modes) with theoretical guarantees on both privacy and accuracy, achieving near-optimal statistical rates that balance the privacy-utility tradeoff.
This paper develops methods for finding density modes (peaks in probability distributions) while guaranteeing differential privacy—a mathematical constraint that limits what can be learned about individual data points. The authors propose DP-GRAMS, which uses noisy gradient ascent on a privately estimated score function, and prove it recovers all modes with near-optimal accuracy.
Juncheng Zhong, Chenghuang Shen, Jianfeng Liu et al.
Decoupling field reconstruction from equation selection—by freezing a learned field representation and then selecting terms via stability-validated weak-form analysis—improves PDE discovery from sparse observations compared to end-to-end neural approaches.
This paper tackles PDE discovery from sparse data by separating field reconstruction from equation selection. The authors develop a freeze-then-select method that first trains a neural adapter to reconstruct the continuous field, then uses stability-validated weak selection to identify the correct differential terms.