Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.
Haojin Deng, Zhiping Lin, Yimin Yang
Monitoring centroid geometry during training can help detect and reduce spurious feature reliance, but attribute information remains partially recoverable—suggesting regularization alone isn't sufficient for complete bias removal.
BiasFlow is a monitoring toolkit that tracks how neural network backbones rely on spurious features (like gender in face recognition) through geometric analysis of feature centroids.
Xuanyu Lu, Fengqing Jiang, Kaiyuan Zheng et al.
Small, fixed adversarial patches can severely degrade world action models across multiple robotic tasks by exploiting how visual encoders process information, highlighting that securing shared visual components is critical for robust robotic control systems.
This paper presents TAPDreamer, an attack method that uses small visual patches to fool world action models—AI systems that predict how robotic environments will change. Unlike previous attacks, TAPDreamer works without accessing the target model, instead using a public encoder to create a single patch that transfers across different tasks and robot policies.
Rubén Manrique, Michelle Castellanos, Jorge Morales et al.
LLMs can sound authoritative about law they don't actually know; current models need expert oversight and source grounding for real legal work, especially outside the US where training data is sparse.
This paper evaluates how well large language models understand Colombian law by testing 15 models on 1,042 expert-validated questions covering ten legal areas. While models score well on multiple-choice questions (up to 90.5%), their free-text legal answers are rarely correct (max 45%), and they often sound confident while being wrong—a dangerous combination for non-experts relying on legal AI.
Pengfei Li, Naufal Suryanto, Sicheng Zhang et al.
Current open-weight LLMs struggle with precise cybersecurity tool use (max 42% accuracy), but fine-tuning with verifiable rewards from this benchmark can make smaller models competitive with much larger ones.
KaliBench is a benchmark for evaluating how well language models can translate security analyst requests into executable commands for Kali Linux tools. It includes 8,504 query-command pairs across 1,642 tools and provides a verification system that checks both whether commands are syntactically correct and whether they actually run successfully, without needing to execute them during training.
Kevin Jiang, Morgane Austern, Edgar Dobriban et al.
You can align generative AI outputs to target distributions by intelligently filtering multiple model queries—no model retraining needed—and this approach is provably optimal for large batches of outputs.
This paper addresses how to align AI-generated outputs with user-specified attribute distributions through post-processing, without modifying the model itself. The authors develop algorithms that select outputs from multiple queries to a generative model, ensuring attributes like gender or age match target distributions.
Ali Holmov, Yiran Huang, Kirill Bykov et al.
LLMs maintain readable and writable internal user models that directly influence safety behavior; different models independently converge on similar user representations, suggesting this is a fundamental property of how language models condition their responses.
This paper introduces Belief Self-Distillation (BSD), a technique to extract and manipulate how LLMs represent their users internally.
Renkai Ma, Ruyuan Wan, Xuan Lu et al.
Building AI agents that users trust requires focusing on operating conditions—cost, oversight, and access controls—not just task performance. Users care deeply about being able to supervise and review agent actions.
This study analyzed 73,000+ Reddit posts about using OpenClaw (an AI agent tool) to understand what values matter to users beyond just task completion.
Parivesh Priye, Yufeng Wang, Haibin Ling et al.
When deploying safety-critical ML systems with selective prediction, finite calibration data is the bottleneck—smart partition selection and error budget reallocation can recover 60% of the theoretical coverage gain, but naive approaches recover almost none.
This paper addresses how to safely deploy selective predictors (models that abstain when uncertain) by certifying they meet precision targets for specific user groups. The key challenge is that with limited calibration data, some groups may not have enough evidence to certify safety.
Hongbo Chen, Li Charlie Xia
You can now rigorously measure and estimate how much your model's performance will degrade when facing distribution shifts, using a unified framework that works across different types of shifts and loss functions.
This paper addresses how machine learning models fail when training and test data distributions differ.
Yakov Pyotr Shkolnikov
Agentic systems that persist across task boundaries need built-in adaptive drives for behavioral regulation, but this same persistence mechanism that enables useful adaptation can also propagate misalignment—requiring new alignment boundaries around state, authority, and constraints rather than...
This paper proposes an 'artificial id'—an internal adaptive drive mechanism for agentic AI systems that operate continuously across task boundaries. Rather than relying on external specifications for when to continue, stop, or change behavior, the system learns to regulate its own actions through differential persistence.
Wonje Jeung, Sangyeon Yoon, Hyesoo Hong et al.
Vision-language reward models for robotics are fragile to paraphrasing—rewording the same goal can flip success/failure judgments on identical robot behavior, a critical flaw for reliable robotic learning systems.
Vision-language models are being used to score robot behavior, but they fail a basic requirement: giving the same score when instructions are paraphrased. This paper introduces ROBORMBENCH, a benchmark of 2,390 real robot trajectories with 21,673 paraphrases, showing that current VLMs flip between calling identical robot actions successful or failed depending on how you word the goal.
Urja Pawar, Rajitha Ramanayake, Nabeel Kemal et al.
LLM explanations correlate poorly with measured factor importance; operators relying on them to understand or oversee model decisions may be misled, requiring additional verification methods.
This paper tests whether LLM explanations actually match their decision-making by checking if cited factors are truly necessary (changing them changes outputs) or sufficient (keeping them preserves outputs).
Junjie Zhang, Hui Liu, Kecheng Chen et al.
RedEvoAgent learns reusable attack skills from past jailbreak attempts, making red-teaming more efficient and interpretable while avoiding the context bloat and retrieval bias of trajectory-based methods.
RedEvoAgent is an automated red-teaming system that tests LLM agents for security vulnerabilities by learning and refining attack strategies. Unlike previous methods that use fixed attacks or store entire attack histories, it distills successful attacks into concise, interpretable skills that evolve through practice—similar to how a human attacker would learn what works.
Yisen Xi
When deploying LLM agents in regulated environments, separate persona (instructions/tone) from execution (work/state) into different trust domains with a governed contract bridge—this lets you evolve agent behavior freely while maintaining execution auditability and data security.
This paper presents Persona-Execution Separation (PES), an architecture pattern for LLM agents in regulated organizations that need to evolve their instructions and tone freely while keeping their work auditable and traceable.
Adriana Watson, Marco Bücheler, Grant Richards
LLMs can help create regulatory compliance documents, but their effectiveness depends heavily on how clearly the regulation defines the required format—strict rules ensure consistency but risk false information, while flexible rules need better prompts to stay complete.
This paper evaluates how well large language models can generate compliance documents required by EU regulations like GDPR and the Ecodesign for Sustainable Products Regulation.
Chengxiao Wang, Enyi Jiang, Xiaojing Liao et al.
You can improve LLM safety without sacrificing utility by conditionally routing safety rules through a learned gate—CLEAR reduces harmful outputs by 98% while maintaining performance on standard benchmarks.
This paper introduces CLEAR, a method that selectively applies safety training to LLMs using a lightweight gate that controls when safety rules activate. Instead of globally applying safety constraints (which hurts performance on normal tasks), CLEAR routes safety adaptations only when needed, reducing harmful outputs while preserving the model's ability to answer legitimate questions accurately.
Taenyun Kim, Edyta Bogucka, Daniele Quercia
Building AI systems by aggregating public moral votes doesn't guarantee fairness; the developers' upstream design choices (feature selection, voter composition, question wording) are actually normative decisions that reshape what the public thinks they're voting for.
This paper reveals that moral AI systems trained on public votes aren't neutral—developers' hidden choices about which features to vote on, who votes, and how questions are worded significantly shape the resulting AI behavior.
Shangao Li, Yao Zhang, Volker Tresp et al.
Don't trust matched evaluation scores for coding agents—they hide failures introduced by command serialization and parsing.
This paper reveals that standard evaluation metrics for LLM coding agents can hide critical failures in command execution. By testing how Bash commands survive serialization and parsing in different system configurations, the authors show that matched scores mask up to 73% of actual failures—failures introduced not by the model but by how its output is processed.