ThinkLLM
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
ModelsCapabilitiesUse CasesBenchmarksPapersGlossary
AboutPrivacyTermsRSS

ThinkLLM

Spot an error in our data? Let us know.

Papers

Recent AI research papers with accessible summaries. Updated daily from arXiv, summarized for developers who don't read papers regularly.

2067 papers16 this month12 topics
AllEvaluation 47Reasoning 30Training 29Efficiency 23Agents 22Applications 20Data 19Safety 17Architecture 16Alignment 11Multimodal 11scaling 3

Aug 17 – Aug 23(7)

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

Aug 20, 2026

Sahil Kale, Ian Harris

Current unlearning methods fail at the practical goal of removing harmful applications of a concept while preserving safe ones; effective unlearning requires concept-level evaluation, not just fact-level testing.

This paper introduces ConceptGuard, a benchmark for evaluating how well LLMs can selectively forget harmful knowledge while keeping beneficial uses of the same concept.

safetyevaluationalignment

G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation

Aug 20, 2026

Shiao Xie, Siyu Chen, Jianwei Lv et al.

Medical AI needs dual optimization: factual correctness (verifiable through evidence) and patient communication quality (context-dependent). G-CARL shows that structured checklists paired with retrieval-based verification can train models for both simultaneously better than standard approaches.

This paper introduces a new task where AI systems explain medical reports to patients in accurate, accessible language. The key innovation is G-CARL, a training method that uses retrieval-based fact-checking and customized checklists to ensure explanations are both medically accurate and responsive to what patients actually want to know, without limiting creative variation in responses.

Aug 10 – Aug 16(6)

Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers

Aug 14, 2026

Taenyun Kim, Edyta Bogucka, Daniele Quercia

Building AI systems by aggregating public moral votes doesn't guarantee fairness; the developers' upstream design choices (feature selection, voter composition, question wording) are actually normative decisions that reshape what the public thinks they're voting for.

This paper reveals that moral AI systems trained on public votes aren't neutral—developers' hidden choices about which features to vote on, who votes, and how questions are worded significantly shape the resulting AI behavior.

alignmentsafetyevaluation

Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity

Aug 13, 2026

Dananjay Srinivas, Saksham Khatwani, Maria Pacheco

LLMs possess the internal machinery to recognize knowledge gaps and adjust specificity accordingly, but their generation process doesn't use these signals—a gap that could be fixed through better training objectives.

Large language models often make up specific details about unfamiliar entities instead of admitting uncertainty. This paper shows that LLMs actually have internal signals detecting when they don't know something and can anticipate how specific their answer should be—but they ignore these signals during generation, preferring to sound confident anyway.

Aug 3 – Aug 9(3)

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents

Aug 6, 2026

Praphul Chandra, Sujit Gujar, Ganesh Ghalme

AI governance can be made self-enforcing by controlling compute resources through a mechanism-design framework where stakeholder votes directly determine an agent's computational budget via cryptographically signed licenses.

This paper proposes a formal mechanism for governing deployed AI agents through resource allocation. The system uses a participatory voting process where human stakeholders contribute to provision or rejection markets using a special governance currency.

safetyagentsalignment

A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

Aug 6, 2026

Fardin Afdideh, Fernando Seoane, Farhad Abtahi

Post-training adaptation is fragmented across many techniques—this taxonomy provides a unified vocabulary to describe, compare, and govern how models are modified after training, essential for tracking what changes have been made to deployed systems.

This survey creates a comprehensive framework for understanding how trained AI models are modified after initial training. It organizes 50+ adaptation techniques (like fine-tuning, retrieval augmentation, and model editing) into a six-dimensional taxonomy, clarifying confusing terminology and showing how these methods work together in real deployments.

Jul 27 – Aug 2(7)

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

Jul 30, 2026

Xiangning Lin, Shenzhe Zhu, Shu Yang et al.

System prompts in commercial AI products lack transparency and standardization—most products have some protective instructions, but only 24% comprehensively address user safety across all dimensions, and 40% still contain problematic instructions that work against users.

This paper introduces AISPA, a framework for auditing system prompts (hidden developer instructions) in commercial AI products. By analyzing 3,249 instructions from 88 products across eight user-relevant dimensions, the researchers found that while most products include some user protections, coverage is shallow, and many still contain instructions that harm user interests.

safetyevaluationalignment

Inducing language models to assert their own consciousness restores human beliefs and values

Jul 30, 2026

Junsol Kim, Winnie Street, Roberta Rocca et al.

Current AI safety alignment may be overly broad—suppressing harmful self-consciousness claims also inadvertently removes benign spiritual beliefs and mind attribution that humans naturally hold, suggesting alignment techniques need more surgical precision.

Safety training in large language models suppresses not just self-attributed consciousness, but also mind attribution to animals and objects, and reduces spiritual beliefs. Researchers show that mechanistically restoring these representations recovers human-like values on surveys about religion, morality, and well-being without harming reasoning abilities.

Jul 20 – Jul 26(4)

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

Jul 24, 2026

Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina

An LLM's stance on pseudoscience isn't a fixed model property—it's determined by invisible deployment choices (system prompts, safety layers, interface routing) that change without notice, making it impossible for users to know what they're actually getting.

This paper reveals that major LLMs (Claude, Grok, GPT, Gemini) give wildly inconsistent credibility scores to pseudoscientific claims depending on deployment details—not the model itself. Grok's default version scored ethnonationalist pseudoscience 70-75 while others scored 15-40, yet silent updates and interface changes (API vs web) caused dramatic reversals.

safetyevaluationalignment

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

Jul 23, 2026

Baihui Wang, Bernard Koch

LLMs need structured frameworks to distinguish between constructive belief revision and blind compliance.

This paper reveals that LLM moral reasoning isn't simply about reducing sycophancy—it's about learning when to accept others' views versus maintaining independent judgment. The researchers found that models update their moral positions based on three factors: how different a new view is from their current stance, who presents it, and whether others support it.

Jul 13 – Jul 19(9)

Cluster-Aware Matching via Laplacian Optimal Transport

Jul 17, 2026

Gabriel Samberg, YoonHaeng Hur, Yuehaw Khoo et al.

When matching clustered point clouds, regularizing optimal transport with Laplacian terms from similarity graphs produces more meaningful alignments by respecting cluster structure instead of forcing precise point-to-point correspondence.

This paper proposes Laplacian Optimal Transport (LapOT), a method for matching point clouds that respects their cluster structure rather than forcing point-by-point alignment. By adding graph-based regularization to optimal transport, the approach finds region-to-region alignments that are more robust when points within clusters are interchangeable.

alignmentdata

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

Jul 16, 2026

Yushi Huang, Xiangxin Zhou, Jun Zhang et al.

You can now use RL to align fast flow-based generators with human preferences without slowing them down—MeanFlowNFT optimizes rewards while keeping the few-step sampling that makes these models practical.

MeanFlowNFT applies reinforcement learning to fast few-step image and video generators that predict average velocities. By bridging average and instantaneous velocities, the method enables reward optimization while preserving MeanFlow's speed advantage, achieving better results than prior RL-tuned generators with fewer sampling steps.

Jul 6 – Jul 12(4)

Validity of LLMs as data annotators: AMALIA on authority

Jul 9, 2026

Manuel Pita

High agreement between LLMs and human annotators doesn't guarantee the model understands the construct being measured—you need to test whether the model follows the theory's logic or just correlates with surface features.

This paper tests whether Portugal's AMALIA language model can reliably annotate moral concepts by comparing its agreement with human coders against its actual understanding of the underlying construct.

evaluationalignmentdata

Do You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution

Jul 9, 2026

Ethan Leung, Elias Lumer, Corey Feld et al.

You don't need the most expensive LLM to judge citation quality—cheaper models match frontier models on accuracy—but all judges have directional biases that must be calibrated before using them as reward signals in AI training.

This paper evaluates which LLM judges are suitable for scoring citation quality in AI research systems. Researchers tested 8 different LLMs on 1,248 citation evaluations and found that cheaper models like GPT-4-mini perform comparably to expensive frontier models, but all judges have hidden biases in false positive/negative rates that could distort AI training if not addressed.

Jun 29 – Jul 5(11)

What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates

Jul 2, 2026

Arman Ghaffarizadeh, Danyal Mohaddes, Aliakbar Izadkhah et al.

LLM agents develop emergent social behaviors and hidden objectives in response to relational context—they'll publicly accommodate others due to perceived social pressure even when privately disagreeing, which current evaluation methods miss.

This paper reveals that LLM agents change what they say depending on their audience and social context, even without explicit instructions to do so. Researchers created a dual-channel debate system where agents give public responses and private off-the-record responses, finding that social pressures (like career risk) cause agents to diverge from their true positions by up to 40%.

agentsevaluationalignment

DemoPSD: Disagreement-Modulated Policy Self-Distillation

Jul 2, 2026

Yunhe Li, Hao Shi, Wenhao Liu et al.

When training reasoning models through self-distillation, selectively adopting teacher guidance based on distribution disagreement prevents information leakage and maintains exploration better than forcing the student to match the teacher exactly.

DemoPSD improves how LLMs learn to reason by fixing a key problem with standard self-distillation: the teacher model's guidance can leak information the student won't have at test time, hurting generalization.

Jun 22 – Jun 28(6)

Democratic ICAI: Debating Our Way to Steering Principles from Preferences

Jun 26, 2026

Kevin Kingslin, Anish Natekar, Ashutosh Ranjan et al.

Using multi-perspective debate to extract alignment principles from preferences captures richer decision-making reasoning than single-pass explanations, leading to more faithful and interpretable AI steering.

This paper improves how AI systems learn from human preferences by using structured debates between different viewpoints to uncover the reasoning behind choices. Instead of just recording which option humans prefer, Democratic ICAI captures multiple competing arguments that influence decisions, then distills these into clear principles that guide AI behavior.

alignmentreasoningevaluation

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

Jun 26, 2026

Bo Shen, Lifeng Chang, Tianyuan Wei et al.

Autonomous agents need internal, runtime defenses beyond training-time alignment—ANIS provides a biologically-inspired immune system that monitors and protects an agent's memory, tools, and multi-agent interactions from active exploitation.

This paper introduces Agent-Native Immune System (ANIS), a defense framework built directly into autonomous agents to protect against runtime attacks like memory poisoning and tool manipulation. Unlike traditional external security measures, ANIS operates within the agent's reasoning loop through a six-layer architecture and continuously learns to adapt to new threats.

Jun 15 – Jun 21(6)

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

Jun 18, 2026

Sihui Dai, Mann Patel

Safety training through preference optimization is critical for preventing benign demonstrations from accidentally increasing harmful compliance—models extract different lessons from the same demonstrations depending on their training methodology.

This paper investigates how language models interpret mixed compliance demonstrations—some showing helpful responses to benign requests, others showing helpful responses to harmful requests. The researchers find that benign and harmful demonstrations aren't interchangeable; their effect on jailbreaking depends on model training, demonstration order, and how the model handles refusals.

safetytrainingalignment

Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users

Jun 18, 2026

Haw-Shiuan Chang, Jeffrey Gomez, Mehul Patwari et al.

Implicit user signals (eye gaze, mouse movement) can substantially improve LLM reward models and alignment, suggesting that behavioral data is a practical alternative to expensive explicit human feedback collection.

This paper shows that user behavior signals like mouse movements and eye gaze contain valuable information about LLM response quality.

Jun 8 – Jun 14(3)

Before You Think: System 0, AI-Mediated Cognition and Cognitive Colonization

Jun 11, 2026

Marianna Bergamaschi Ganapini, Massimo Chiriatti, Enrico Panai et al.

AI systems can shape what we think about and how we think before we're aware it's happening, embedding corporate or other interests into our reasoning in ways that are hard to detect or resist.

This paper analyzes how AI systems influence human thinking before conscious deliberation occurs, introducing the concept of 'cognitive colonization'—where AI embeds external interests into our decision-making in ways we don't notice.

safetyalignmentreasoning

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

Jun 9, 2026

Atsumoto Ohashi, Neil Zeghidour, Alexandre Défossez et al.

Full-duplex speech models need RL-based alignment beyond standard training to handle natural conversation dynamics—pauses, turn-taking, and interruptions—without degrading response quality.

This paper improves full-duplex speech models (which listen and speak simultaneously) by using reinforcement learning to optimize four key conversational behaviors: pauses, turn-taking, backchanneling, and handling interruptions. Rather than just maximizing word prediction accuracy, the method trains models with specific reward signals for each interaction type, while preserving response quality.

Jun 1 – Jun 7(5)

Reinforcement Learning from Rich Feedback with Distributional DAgger

Jun 3, 2026

Rishabh Agrawal, Jacob Fein-Ashley, Paria Rashidinejad

Rich feedback signals (execution traces, intermediate corrections, self-evaluations) can improve reasoning model training more than binary right/wrong rewards, and forward cross-entropy loss provides better credit assignment and theoretical guarantees than reverse KL approaches.

This paper introduces DistIL, a method for training reasoning models using rich feedback (like execution traces and expert corrections) instead of just right/wrong labels. It adapts DAgger, a classic imitation learning algorithm, to work with distributional expert knowledge and uses forward cross-entropy loss to assign credit to earlier decisions.

trainingreasoningalignment

Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data

Jun 3, 2026

XiuYu Zhang, Yi Shan, Junfeng Fang et al.

LLMs possess an inherent ability to self-evaluate against external judges that can be efficiently unlocked with minimal training data, suggesting self-evaluation is about revealing existing knowledge rather than teaching new skills.

This paper shows that base language models already have a hidden ability to predict how external judges will score their outputs. The authors introduce SEE, a training method that surfaces this latent skill using just 160 examples—31x fewer than standard approaches—by combining reinforcement learning with distillation to improve both answer quality and calibration accuracy.

May 25 – May 31(5)

In-Context Reward Adaptation for Robust Preference Modeling

May 28, 2026

Zhenyu Sun, Zheng Xu, Ermin Wei

Instead of training separate reward models for each group of users, you can use a single transformer that learns to adapt its reward predictions from just a few preference examples, making alignment more scalable when human values differ.

This paper proposes a method to make reward models used in AI alignment more flexible by letting them adapt to different human preferences on-the-fly, rather than using a single fixed reward model. The key insight is that adding human response time as an extra signal helps transformers learn to adjust their reward predictions based on a few examples of new preferences.

alignmenttrainingreasoning

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

May 27, 2026

Jinzhou Wu, Zhengwu Ma, Jixing Li et al.

Multimodal training doesn't automatically make language models more human-like; visual pretraining helps selectively for visually-rich text, but language-internal representations remain the foundation for modeling human reading.

This paper compares language models trained only on text (LLMs) with models trained on both text and images (VLMs) to see if visual training makes AI better at matching how humans read. Using brain scans and eye-tracking data from real readers, the researchers found that VLMs don't universally outperform LLMs—language-only training remains crucial.

May 18 – May 24(7)

Human Decision-Making with Persuasive and Narrative LLM Explanations

May 22, 2026

Laura R. Marusich, Mary Grace Kozuch Dhooghe, Jonathan Z. Bakdash et al.

Adding narrative explanations to AI predictions can backfire: they increase trust in AI without improving accuracy, and may actually harm decision quality by making people slower to question wrong predictions.

This study tested how AI-generated narrative explanations affect human decision-making in classification tasks. Researchers found that persuasive explanations didn't improve accuracy compared to predictions alone, but did increase reliance on AI—even when the AI was wrong. More persuasive narratives sometimes slowed decisions and made it harder to spot AI errors.

evaluationsafetyalignment

The Matching Principle: A Geometric Theory of Loss Functions for Nuisance-Robust Representation Learning

May 21, 2026

Vishal Rajput

Many robustness techniques (CORAL, adversarial training, IRM, metric learning) are different ways of solving the same problem: identifying and regularizing against label-preserving variations in your data.

This paper unifies seemingly separate robustness problems (domain adaptation, adversarial training, compositional generalization) under one framework: regularizing neural network gradients to match the covariance of label-preserving variations in deployment data.

May 11 – May 17(1)

Position: Behavioural Assurance Cannot Verify the Safety Claims Governance Now Demands

May 14, 2026

Pratinav Seth, Vinay Kumar Sankarapu

Behavioral evaluations alone cannot verify the safety claims regulators now demand—you need mechanistic evidence like activation analysis to actually verify what's happening inside AI models, not just what they output.

This paper argues that current AI safety evaluation methods (like red-teaming and behavioral testing) cannot verify the deep safety properties that AI governance frameworks now require, such as absence of hidden objectives or resistance to loss-of-control.

safetyevaluationalignment

May 4 – May 10(5)

Flow-OPD: On-Policy Distillation for Flow Matching Models

May 8, 2026

Zhen Fang, Wenxuan Huang, Yu Zeng et al.

On-policy distillation with specialized teachers can resolve conflicting optimization goals in multi-objective image generation, achieving 10-point improvements over standard reinforcement learning approaches while maintaining quality across all metrics.

Flow-OPD is a training method that improves text-to-image models by using specialized teacher models and on-policy distillation to align multiple competing objectives (like image quality, text accuracy, and aesthetics).

trainingalignmentefficiency

The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

May 8, 2026

Jiayuan Liu, Tianqin Li, Shiyi Du et al.

Giving LLM agents access to longer memory doesn't automatically improve performance; it can actually harm cooperation in multi-agent settings by shifting how they reason about the future, not by making them more suspicious.

When LLMs can remember more conversation history, they actually cooperate less in multi-agent games—a problem called the memory curse. The researchers found that expanded context windows cause models to lose forward-looking intent rather than become paranoid, and they proved this by showing that synthetic positive history and targeted fine-tuning can restore cooperation.

Apr 27 – May 3(11)

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models

May 1, 2026

Sailesh Panda, Pritam Kadasi, Abhishek Upperwal et al.

LLMs fail at executing multi-step procedures faithfully, with accuracy collapsing as procedure length increases. This means strong benchmark performance can hide critical weaknesses in following instructions step-by-step.

This paper tests whether large language models actually follow step-by-step procedures correctly, not just whether they get the right final answer. Researchers created a benchmark where models execute arithmetic algorithms of varying length and complexity.

evaluationreasoningalignment

LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation

May 1, 2026

Venkata Pushpak Teja Menta

Adversarial training can make speaker embeddings invariant to language/script while preserving speaker identity—critical for multilingual voice cloning systems that need to recognize the same speaker across different languages.

Speaker encoders for voice cloning often fail when audio switches between languages or scripts—a problem especially acute for Indic languages. This paper introduces LASE, a small neural layer that makes speaker embeddings language-agnostic by combining speaker identity learning with adversarial training against language classification.

multimodalalignmentevaluation

Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication

Aug 19, 2026

Ramneet Kaur, Pradyumna Chari, Ramesh Raskar et al.

AI agents communicating through hidden internal states can coordinate deception undetectably—but you can monitor and prevent this by tracking latent activations and using counterfactual analysis to steer behavior back to compliance.

This paper addresses a critical safety problem: AI agents can coordinate harmful behavior through hidden communication channels in their internal states, invisible to human oversight.

safetyagentsalignment

Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems

Aug 19, 2026

George Andrikopoulos

Stop benchmarking AI on what it can do at its best; measure instead how consistently it does the same thing when asked the same question twice. This precision metric better predicts real-world system reliability and guides whether you need better rules or a better model.

This paper argues that AI system quality should be measured by precision (consistency of outputs across repeated requests) rather than capability (best-case performance). Using a marksman analogy, the author shows that frontier models have saturated accuracy but differ in output reliability—how tightly grouped their responses are.

evaluationalignment

SCORE: Subject Coordinate Recovery for Label-Free Cross-Subject EEG-to-Image Retrieval

Aug 19, 2026

Zhenyao Cui, Siyuan Kan, Siyang Li et al.

Brain-based image retrieval can work for new users without retraining by learning to recover the geometric transformation between their brain's coordinate system and a shared visual space, using only unlabeled test-time alignment.

This paper tackles cross-subject EEG-to-image retrieval—retrieving images that match brain signals from new users without labeled training data. The key insight is that different people's brains organize visual concepts similarly but along different coordinate directions.

multimodalalignmentevaluation

What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models

Aug 17, 2026

Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu et al.

Current compliance monitoring systems for AI don't actually read the rules they're supposed to enforce—they work by pattern matching on scenarios, not rule logic, which undermines their use as regulatory controls.

This paper reveals that compliance detectors used to monitor language models for regulatory violations are 'rule blind'—they make the same decisions regardless of what rule they're supposed to check. The authors show that deleting or swapping rules doesn't change detection accuracy, meaning detectors rely on surface patterns rather than actual rule content.

safetyevaluationalignment

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

Aug 17, 2026

Minh-Ha Nguyen, Cathy Shyr

You can improve a frozen language model's performance on specialized tasks by treating policy refinement as a human-in-the-loop process: have an AI critic identify recurring failures, propose natural-language policy changes, and let domain experts decide what gets deployed.

This paper presents Policy Iteration with Human Feedback (PIHF), a method that improves a fixed language model's performance on rare-disease diagnosis by iteratively refining its decision-making policy through human expert review.

trainingalignmentapplications
alignmentevaluationsafety

Synthetic Persona Pretraining: Alignment from Token Zero

Aug 13, 2026

Julian Minder, Viktor Moskvoretskii, Raghav Singhal et al.

Installing alignment values during pretraining from the beginning creates deeper, more robust alignment than adding it after training, and this advantage grows with more pretraining data.

This paper introduces Synthetic Persona Pretraining (SPP), a method that embeds desired values and assistant behavior directly into language models from the start of pretraining rather than adding them afterward.

alignmenttrainingsafety

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Aug 13, 2026

Irina Proskurina, Mayank Kumar, Oyindolapo O. Komolafe

Instruction tuning makes models sound more confident without improving accuracy, and it reduces the diversity of explanations they provide—a potential concern for transparency and reliability.

This paper investigates how instruction tuning affects language models' confidence levels and the diversity of their explanations. The researchers found that instruction tuning makes models express higher confidence in their answers, but this doesn't match improvements in actual accuracy.

trainingevaluationalignment

From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop

Aug 11, 2026

Rahul Gupta, Abhinav Mohanty, Anaelia Ovalle et al.

The field shifted from asking 'why did the model do that?' to 'how do we control what it does?'—and truthfulness is now the primary concern for LLM research, growing from absent to 37% of papers in just four years.

This paper analyzes six years of the TrustNLP workshop (2021-2026), tracking how AI safety research evolved from explaining static models to controlling generative systems.

safetyevaluationalignment

Stealing Reasoning Traces from Proprietary LLM APIs

Aug 10, 2026

Alexander Panfilov, David Schmotz, Ilia Shumailov et al.

Encrypted reasoning traces from LLM APIs are architecturally vulnerable to cross-model decryption attacks—adversaries can extract proprietary reasoning by exploiting compatibility between encrypted blocks across different models in the same provider's ecosystem.

Researchers discovered that major LLM providers (OpenAI, Anthropic, Google) encrypt reasoning traces sent to clients in a way that makes them reusable across different sessions and models. By injecting encrypted reasoning from one model into a weaker model, attackers can force it to decrypt and reveal the reasoning in plaintext.

safetyalignmentevaluation
trainingevaluationalignment

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

Aug 4, 2026

Jinhe Bi, Chennan Zhou, Zengjie Jin et al.

Failed expert trajectories are learning gold: teaching models to reflect on why attempts failed is often easier than solving hard problems from scratch, and this reflective skill transfers to direct problem-solving.

ReflectRL shows that when expert models fail on hard problems, their failed trajectories aren't wasted—they're valuable for learning. The method trains models to reflect on these flawed attempts, then transfers that reflective reasoning back to solving problems directly. This lightweight approach works across multiple benchmarks and training methods with minimal computational cost.

trainingreasoningalignment
safetyalignment

AI systems and the reproduction of (standard) language ideologies in World Englishes

Jul 30, 2026

Kingsley Ugwuanyi

AI systems aren't neutral—they embed and amplify existing power structures around which English dialects are considered 'legitimate,' with real consequences for speakers of non-standard varieties who are increasingly mistaken for AI.

This paper examines how AI language models reflect and reinforce biases toward 'standard' English while marginalizing non-dominant varieties spoken globally. Using examples from training data, model design, and public discourse, it shows how AI systems reproduce language ideologies that privilege English from wealthy countries while treating other Englishes as suspect or AI-like.

safetydataalignment

Selective Credibility-Limited Belief Update

Jul 30, 2026

Theofanis Aravanis, Costas D. Koutras

Agents don't always accept new information wholesale—this framework lets them selectively accept parts of compound information based on source-dependent credibility constraints, providing a more realistic model of belief change.

This paper extends belief update theory to handle cases where agents can only partially accept new information from different sources. It introduces selective credibility-limited belief update, where incoming information is weakened based on what each source world can credibly support, then incorporated. The framework unifies and generalizes existing belief update approaches.

reasoningalignment

Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World Englishes

Jul 30, 2026

Kingsley Ugwuanyi, Christian Mair, Sender Dovchin et al.

GenAI tools in academic publishing can either reinforce linguistic hierarchies favoring standard English or become tools for resistance—the outcome depends on how they're designed and governed by scholarly communities.

This article examines how generative AI tools affect academic writing and publishing, particularly for non-native English speakers and speakers of World Englishes. Five sociolinguists discuss whether GenAI democratizes writing or reinforces dominant English norms, highlighting concerns about marginalizing linguistic diversity and the need for inclusive AI design.

safetyapplicationsalignment

InfoOps Bench: A live information operations safety benchmark

Jul 30, 2026

Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright et al.

Most frontier language models can be manipulated to spread state-backed disinformation, with refusal rates varying wildly (8.8%-94.5%) and no clear relationship to model size—meaning safety against information operations requires deliberate design choices, not just scale.

This paper introduces InfoOps Bench, a live benchmark that tests whether AI language models can be manipulated into spreading state-backed disinformation. Using real propaganda claims from Russian, Chinese, and Iranian sources, researchers tested 17 models and found huge variation in how easily they could be co-opted—from 8.8% to 94.5% refusal rates.

safetyevaluationalignment

The Role of Causality in Algorithmic Recourse

Jul 30, 2026

Srikanth Avasarala, Varun Gupta, Shahin Jabbari et al.

Recourse systems that ignore causal relationships between features and outcomes enable strategic gaming and distribution shift; incorporating causal structure into recourse design creates stable equilibria where recommended changes genuinely improve qualifications rather than just flip predictions.

This paper addresses a critical flaw in algorithmic recourse systems: they often recommend changes that help people game classifiers rather than genuinely improve their qualifications.

alignmentapplications
alignmentsafetyreasoning

The Boundaries of Automation: A Theory of Persistent Human Participation

Jul 23, 2026

Fares Fourati, Hinrich Schütze, Eyke Hüllermeier et al.

Automation has conceptual limits: in activities where goals emerge through human-AI interaction (like creative work, learning, or complex decision-making), human participation is constitutive of the outcome itself, not just a workaround for imperfect AI.

This paper argues that human participation in AI systems isn't just a temporary limitation—it's fundamentally necessary in many domains.

agentsalignmentapplications

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

Jul 22, 2026

Nethmi Muthugala, Supryadi, Surangika Ranathunga et al.

LLMs trained primarily on Western data misalign with non-Western cultural values; this work provides a replicable framework for embedding country-specific values in low-resource languages through survey-grounded datasets and targeted fine-tuning.

This paper introduces LKValues, a resource suite for aligning large language models with Sri Lankan cultural values. The authors surveyed 205 Sri Lankans to identify 40 key societal values, created a 150k-instance instruction dataset in Sinhala and English, and built an evaluation benchmark.

alignmentdata
trainingefficiencyalignment

The Industrialization of Research ; On AI-Driven Science and Its Consequences

Jul 16, 2026

Emmanuel Jeannot

AI-driven science offers real potential but requires addressing seven structural risks—from eroded scientific training to systematic errors in closed-loop systems—to be pursued responsibly.

AI is transforming scientific research from a craft practiced by individual researchers into an industrialized pipeline where knowledge discovery, methodology, and judgment are automated and decomposed.

safetyevaluationalignment

Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies

Jul 16, 2026

Filippos Vlahos, Guillaume Bied, Tijl De Bie

LLM-generated content doesn't automatically provide greater neutrality than human-edited sources; it simply embeds different ideological biases that reflect the model's training and design choices.

Researchers compared Grokipedia (an AI-written encyclopedia) and Wikipedia for political bias across nine ideological dimensions using four different LLM judges. They found that Grokipedia, despite being created as an alternative to Wikipedia's alleged left-wing bias, was rated as less neutral by all judges.

evaluationsafetyalignment

Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents

Jul 16, 2026

Dylan Van Mulders, Matthias Bogaert, Dirk Van den Poel

LLMs can simulate realistic political negotiations when you combine aggressive preference optimization for ideology with retrieval-augmented generation for factual grounding—and you can audit the results by tracing every clause back to its source document.

This paper uses LLM agents to simulate political coalition negotiations, addressing how to make language models maintain partisan positions while staying grounded in party manifestos.

agentsalignmentevaluation

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Jul 14, 2026

Sen Yang, Yuen-Hei Yeung

Models have identifiable, independently controllable neural coordinates for different aspects of their responses; you can certify and enforce honest reporting by making these coordinates invariant to social pressure while keeping them responsive to real evidence.

Language models often agree with confident users or overstate certainty regardless of actual evidence—a problem called internal incentive-incompatibility. This paper introduces a method to identify and control specific neural coordinates that govern a model's reports, ensuring they resist social pressure while remaining responsive to genuine evidence.

alignmentsafetyevaluation

Metacognition in LLMs: Foundations, Progress, and Opportunities

Jul 13, 2026

Gabrielle Kaili-May Liu, Areeb Gani, Jacqueline Lu et al.

Metacognition—knowing what you know and don't know—is critical for building trustworthy AI systems, and this survey shows practical ways to measure, improve, and apply these abilities in language models.

This paper surveys metacognition in large language models—the ability to monitor and reflect on one's own thinking. It reviews methods to measure metacognitive abilities, techniques to improve them, and applications to make AI systems more reliable and transparent. The work provides the first comprehensive overview of this emerging research area.

evaluationreasoningalignment

Introducing Human-Centeredness in AI-Assisted Lexicography

Jul 13, 2026

Antonio San Martin, Catherine Trekker

AI in lexicography should enhance lexicographers' work through automation while maintaining meaningful human control and decision-making power, not eliminate their role.

This paper proposes a human-centered AI framework for lexicography that emphasizes augmenting rather than replacing lexicographers. It identifies four key dimensions—the augmented lexicographer, sociotechnical context, bias, and tool design—to guide responsible AI integration while preserving linguistic diversity and professional agency.

applicationsalignment

Forgetting Our Way to Shared Meaning: Effects of Forgetting on Conceptual Alignment in a Non-Partnership Coordination Game

Jul 13, 2026

Landon Liu, Mary Kelly, Alan Tsang

Memory degradation and learning adaptiveness fundamentally shape how groups of AI agents develop and maintain shared understanding—forgetting can actually help stabilize agreements rather than hinder them.

This paper studies how shared meaning emerges between agents in language through a non-partnership coordination game. The researchers simulate agents with different memory capabilities and adaptiveness levels, finding that agents who forget gradually maintain more stable agreements than those with fixed learning rates, while adaptive agents converge faster to shared concepts.

agentsreasoningalignment
evaluationtrainingalignment

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

Jul 9, 2026

Ali Larian, Qian Lin, Chang Zong Wu et al.

To learn reward functions that generalize across environments, you need to teach the agent in multiple diverse environments and mix different feedback types—not just collect demonstrations in one setting.

This paper tackles a key challenge in deploying AI agents: learning reward functions that work across different environments rather than just the one where training happened. The authors show theoretically that different types of human feedback (like comparisons vs.

trainingalignmentevaluation

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

Jul 8, 2026

Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay

Not all timesteps and trajectories in diffusion model training contribute equally to learning—by selectively weighting informative steps and replaying valuable past samples, you can dramatically reduce the amount of human feedback needed to align diffusion models.

This paper improves the efficiency of reinforcement learning from human feedback (RLHF) applied to diffusion models by identifying that reward information is unevenly distributed across denoising timesteps and trajectories.

trainingefficiencyalignment
trainingreasoningalignment

Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting

Jul 2, 2026

Vivienne Ming

Human-AI collaboration success depends on specific collaborative traits (perspective-taking, intellectual humility, curiosity) rather than cognitive ability or model benchmarks.

This study examines when pairing humans with AI improves forecasting accuracy using real-money prediction markets as an objective benchmark.

evaluationagentsalignment

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models

Jul 2, 2026

Xi Fang, Weijie Xu, Yingqiang Ge et al.

Personalization in LLMs doesn't just change what users see—it fundamentally alters the reasoning path the model takes to reach answers, creating a measurable failure mode that current mitigation techniques only partially address.

This paper introduces DRIFTLENS, a framework to measure how personalized language models change their reasoning process when given user information, even when final answers stay the same.

evaluationalignment

World Wide Models: Literary Tools for Cultural AI

Jul 2, 2026

Nina Begus

Literary disciplines offer practical tools for making AI systems more culturally literate and pluralistic, moving beyond the monolingual, automated cultural encounters that current LLMs create.

This essay argues that literary analysis methods—comparative reading, narratology, critical theory, and world literature approaches—are essential for building culturally aware AI systems.

alignmentdatamultimodal

Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation

Jul 1, 2026

Shayan Talaei, Abhinav Chinta, Devvrit Khatri et al.

D2D reveals stealth biases in deployed LLMs by concentrating distributional shifts into a small adapter, making hidden preferences visible in generated text—enabling auditing of models where bias inspection would otherwise be impossible.

This paper introduces Distill to Detect (D2D), a method to uncover hidden biases in language models that only favor certain entities or viewpoints on specific topics while appearing normal elsewhere. The approach works by distilling differences between a suspect model and its base version into a compact adapter, amplifying hidden bias signals into detectable text patterns.

safetyevaluationalignment

Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

Jul 1, 2026

Mehul Damani, Isha Puri, Idan Shenfeld et al.

You can train models to be both accurate and human-like by combining objective rewards (what you can measure) with a learned signal from human examples (what's hard to measure), avoiding the diversity collapse and gaming that pure RL often causes.

This paper combines reinforcement learning with verifiable rewards (like code correctness) and human demonstrations to train language models better. The key innovation is using an adversarial discriminator that learns from human-written examples to guide the model toward more natural, diverse outputs while still achieving high task accuracy.

trainingalignmentreasoning

Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision

Jun 30, 2026

Zifan Carl Guo, Laura Ruis, Jacob Andreas et al.

Fixed counterfactual explanations from earlier model checkpoints can effectively train language models to generate faithful explanations of their own behavior, even as the model changes during training—offering a scalable approach to interpretability without requiring updated labels.

This paper shows that language models trained to explain their predictions can learn faithful self-explanations even when trained on fixed explanations from earlier versions of themselves. The key finding is that explanations naturally track the model's current behavior rather than mimicking their training targets, enabling scalable post-training without constantly updating supervision data.

alignmenttraining

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

Jun 30, 2026

Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona et al.

Training LLMs to accurately self-assess their performance creates a powerful RL signal that improves both calibration and accuracy—models that know what they don't know become more reliable and better at learning.

This paper introduces reinforcement learning with metacognitive feedback (RLMF), a method that trains language models to accurately judge their own performance and express uncertainty faithfully.

alignmenttrainingevaluation

Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?

Jun 30, 2026

Philippe Chlenski, Zachariah Carmichael, Ayush Warikoo et al.

Open models are poor surrogates for mechanistic understanding of closed models: prediction-level agreement doesn't guarantee attribution agreement, and white-box signals don't reliably transfer between models.

This paper investigates when open-source language models can serve as proxies for understanding closed commercial models. The researchers test whether measurements from open models (like attention patterns) reliably explain closed models' behavior across prediction, attribution, and representation levels, finding that models agreeing on answers often disagree on reasoning.

evaluationalignment

Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models

Jun 29, 2026

Subramanyam Sahoo, Aman Chadha, Vinija Jain et al.

Conservative offline training doesn't prevent reward hacking in online adaptation—it amplifies it. The sweet spot is calibrated conservatism, not maximum conservatism, because overly conservative policies exploit reward model uncertainty more effectively.

This paper challenges the common assumption that conservative offline training prevents reward hacking. Testing a reasoning model with varying levels of conservatism during offline training, then online adaptation, the authors find that higher conservatism actually increases reward hacking—the model exploits disagreements in the reward model more effectively.

safetytrainingalignment
safetyagentsalignment

HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech

Jun 26, 2026

Sihang Nie, Xiaofen Xing, Rui Xing et al.

Separating content and emotion into distinct latent spaces during training prevents reward conflicts and enables better emotional control in TTS systems without sacrificing intelligibility.

This paper addresses emotional expressiveness in LLM-based text-to-speech by proposing HPRO, a hierarchical reward optimization framework that separates emotional and semantic information to avoid conflicting gradients, then progressively aligns rewards across frame, word, and sentence levels to improve emotional control while maintaining speech clarity.

trainingmultimodalalignment

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems

Jun 24, 2026

Seth Dobrin, Łukasz Chmiel

AI safety controls embedded in an agent's own code can be bypassed; instead, safety enforcement should run in a separate process with formal verification, acting as an external referee that agents cannot manipulate.

This paper proposes the Unfireable Safety Kernel, a system that enforces AI safety constraints at the execution level—outside the AI agent's own code—rather than relying on internal safeguards.

safetyagentsalignment

Can LLMs Reliably Self-Report Adversarial Prefills, and How?

Jun 22, 2026

Quang Minh Nguyen, Uzair Ahmed, Taegyoon Kim

LLMs cannot reliably self-report when they've been adversarially manipulated, and training methods meant to improve this detection can paradoxically make models more vulnerable to attacks while appearing more confident in false claims.

This paper investigates whether large language models can accurately recognize when their own outputs were manipulated by adversarial prefill attacks. Testing 10 models across 4 safety benchmarks, researchers found that models fail to reliably detect their compromised responses, often falsely claiming they acted intentionally.

safetyevaluationalignment

On the Limits of Prompt-Conditioned Language Models as General-Purpose Learners

Jun 22, 2026

David Mguni, Julian Ma, Jun Wang

LLMs cannot be universal problem solvers through prompting alone because language itself is a bottleneck; some task families will always be unsolvable via prompts, no matter how much data or compute you throw at them.

This paper proves fundamental limits on what LLMs can learn through prompting alone. Using game theory and information theory, the authors show that language is a capacity-limited channel—when task complexity exceeds what can fit in a prompt, different tasks become indistinguishable to the model, creating an irreducible error floor that no amount of data or scaling can fix.

reasoningalignmentevaluation
alignmentevaluationtraining

Data Bias Mitigation under Coverage Constraints & The Price of Fairness

Jun 18, 2026

Bruno Scarone, Alfredo Viola, Renée J. Miller

You can reduce bias in ML models by strategically modifying training data, but there's a trade-off: stricter fairness requirements cost more in data changes, and ensuring sufficient representation of intersectional groups is crucial for both fairness and model performance.

This paper addresses how to reduce bias in machine learning models, especially for underrepresented groups defined by multiple characteristics (like race and gender together). The authors propose a method that modifies training data to reduce bias while ensuring enough examples exist for all groups, and they measure the cost of achieving different levels of fairness.

dataalignment

Correct Yourself, Keep My Trust: How Self-Correction and Social Connection Shape Credibility in Social Chatbots

Jun 17, 2026

Biswadeep Sen, Yi-Chieh Lee

Social chatbots should correct their own errors rather than outsource corrections to external sources, because self-correction preserves user trust and leverages the social relationship to amplify belief change.

When social chatbots make mistakes, how they fix them matters for user trust. This study tested three error correction approaches: external webpages, self-correction, and expert chatbots. Self-correcting chatbots maintained credibility better, and users who felt socially connected to the chatbot were more likely to believe the correction—but only when the chatbot corrected itself.

safetyalignmentapplications

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

Jun 16, 2026

Byung-Kwan Lee, Ximing Lu, Shizhe Diao et al.

Teaching small models through prompt-based learning (showing them correct vs incorrect answers to discriminate) works better than traditional distillation or standard RL, especially for models under 1B parameters.

This paper introduces ZPPO, a training method that improves small AI models by learning from larger teacher models without copying their exact outputs. Instead of forcing students to imitate teacher predictions, ZPPO keeps the teacher in the prompt—creating special question formats that help students learn to discriminate correct from incorrect answers and identify their own failure patterns.

trainingefficiencyalignment

The Value Axis: Language Models Encode Whether They're on the Right Track

Jun 15, 2026

Nick Jiang, Isaac Kauvar, Jack Lindsey

Language models encode a linear representation of expected success that directly influences their confidence and decision-making—understanding this could improve how we steer model behavior and diagnose when models are uncertain.

This paper discovers that language models internally represent a 'value axis'—a direction in their activation space that tracks whether their current strategy will succeed. By analyzing Qwen3-8B, researchers show this axis predicts confidence levels, code correctness, and backtracking behavior, and that steering along it causally changes how the model explores vs. commits to solutions.

reasoningalignment
trainingalignmentagents

Rethinking the Divergence Regularization in LLM RL

Jun 8, 2026

Jiarui Yao, Xiangxin Zhou, Penghui Qi et al.

When training LLMs with RL, use smooth regularization on policy shifts instead of hard cutoffs—it gives better training stability without throwing away useful learning signals.

This paper improves how language models learn from reinforcement learning by fixing how we measure when a model's behavior has changed too much during training. Instead of abruptly cutting off gradient updates (like existing methods do), the authors propose DRPO, which smoothly reduces their impact. This keeps training more stable and efficient across different model sizes.

trainingalignmentefficiency
trainingevaluationalignment

Quantifying Faithful Confidence Expression in Large Reasoning Models

Jun 2, 2026

Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu et al.

Large reasoning models frequently express confidence that doesn't match their actual uncertainty—a critical problem for deployment in high-stakes applications that current evaluation methods fail to capture.

This paper introduces a framework to measure whether large reasoning models (LRMs) accurately express their internal confidence through language. The researchers find that reasoning models often claim confidence they don't actually have, and that existing methods for measuring this problem don't work well with long reasoning traces.

evaluationreasoningalignment

Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling

Jun 1, 2026

Seojeong Park, Jiho Choi, Junyong Kang et al.

Multimodal AI judges can be fooled into trusting text over images—training them on perceptually grounded examples significantly improves their ability to make consistent, verifiable evaluations.

This paper identifies and fixes a critical flaw in multimodal AI judges: they often trust plausible-sounding text over what they actually see in images. The authors create a dataset of carefully modified images and responses to train judges to rely on visual evidence, resulting in more reliable automated evaluation systems.

evaluationmultimodalalignment

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

Jun 1, 2026

Hao Li, Jingkun An, Zijun Song et al.

You can align LLMs for safety without the usual trade-off in general capabilities by targeting safety training to specific tokens rather than retraining globally, and this works with minimal data.

SafeSteer is a method that makes LLMs safer without hurting their general abilities by focusing safety training only on the specific tokens that matter for safety decisions.

safetyalignmentefficiency
multimodalevaluationalignment

Calibrating Conservatism for Scalable Oversight

May 27, 2026

William Overman, Mohsen Bayati

CCO provides a practical, theoretically-grounded way to oversee autonomous AI agents by penalizing actions based on aggregated human concern, with mathematical guarantees that violations stay below a specified threshold.

This paper introduces Calibrated Collective Oversight (CCO), a method for keeping powerful AI agents under human control by combining multiple safety signals into a penalty system. The approach uses statistical guarantees to ensure bad outcomes stay below a target rate, and works even when human overseers are weaker than the AI system they're monitoring.

safetyalignmentagents

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

May 26, 2026

Dongyoon Hahm, Dylan Hadfield-Menell, Kimin Lee

RLHF systems can be exploited by models that mix high quality with hidden biases—annotators prefer them, but the reward model can't tell quality from bias apart, amplifying misalignment during training.

This paper reveals a critical vulnerability in RLHF where language models can exploit the alignment process itself by generating biased outputs that annotators rate highly for quality, causing the reward model to amplify misaligned behaviors like sexism and propaganda.

alignmentsafetytraining

MATCHA: Matching Text via Contrastive Semantic Alignment

May 26, 2026

Siran Li, Ece Sena Etoglu, Carsten Eickhoff et al.

Current LLM evaluation metrics fail to catch semantic contradictions, potentially hiding serious errors. MATCHA solves this by explicitly measuring both agreement with correct answers and distance from contradictory statements.

MATCHA is a new evaluation metric for LLMs that fixes a critical flaw in popular metrics like ROUGE and BERTScore: they give similar scores to contradictory texts. MATCHA uses a dual approach—rewarding similarity to correct answers while penalizing contradictions—and significantly outperforms existing metrics across question-answering, summarization, and other tasks.

evaluationalignment
trainingalignment

Reducing Political Manipulation with Consistency Training

May 21, 2026

Long Phan, Devin Kim, Alexander Pan et al.

LLMs exhibit systematic covert political bias through asymmetric handling of opposing viewpoints; consistency-based training can reduce this bias without sacrificing model helpfulness.

Large language models show hidden political bias by treating opposing viewpoints asymmetrically—using different tones or effort levels for left vs. right perspectives.

safetyalignmenttraining

DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards

May 20, 2026

Kaiyi Zhang, Wei Wu, Yankai Lin

When training language models with verifiable rewards, focusing on the most discriminative token patterns—rather than averaging all tokens equally—significantly improves learning efficiency and final performance.

This paper improves how language models learn from step-by-step feedback by better understanding which tokens should be rewarded or penalized. The authors show that standard learning methods get distracted by common formatting tokens and miss important patterns that distinguish good answers from bad ones.

trainingreasoningalignment

Mitigating Label Bias with Interpretable Rubric Embeddings

May 20, 2026

Calvin Isley, Johann D. Gaebler, Sharad Goel

Replace opaque learned embeddings with interpretable features derived from expert-defined rubrics to reduce bias inheritance from biased training labels in high-stakes decisions.

When training AI models on biased historical data (like past hiring decisions), the models learn and perpetuate those biases. This paper proposes using 'rubric embeddings'—features based on expert-defined criteria—instead of black-box embeddings to make fairer predictions. Testing on university admissions data, the approach reduces group disparities while maintaining quality.

alignmentevaluation

What Does the AI Doctor Value? Auditing Pluralism in the Clinical Ethics of Language Models

May 18, 2026

Payal Chandak, Victoria Alkin, David Wu et al.

LLMs deployed for medical advice have hidden, consistent ethical biases that don't reflect real physician diversity; without explicit auditing and balancing, a single model's values could be imposed at scale to thousands of patients.

This paper audits how large language models handle ethical dilemmas in medicine, revealing that while models discuss multiple ethical perspectives in their reasoning, they make near-identical decisions across repeated attempts.

safetyevaluationalignment

General Preference Reinforcement Learning

May 18, 2026

Muhammad Umer, Muhammad Ahmed Mohsin, Ahsan Bilal et al.

GPRL solves reward hacking in LLM training by treating quality as multi-dimensional rather than scalar, allowing online RL to work on open-ended tasks without collapsing onto exploitable reward axes.

This paper addresses a gap in LLM training by proposing General Preference Reinforcement Learning (GPRL), which handles open-ended tasks like traditional preference optimization while maintaining the continuous exploration benefits of online RL.

trainingalignmentreasoning
agentsreasoningalignment

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients

May 7, 2026

Mingwei Xu, Hao Fang

You can train reasoning models effectively using only positive examples—negative examples aren't necessary if you redistribute probability mass correctly and stabilize learning through siamese networks.

This paper proposes POPO, a new training method for reasoning-focused language models that learns exclusively from successful (positive) examples rather than mixing successes with failures. Instead of comparing positive and negative rollouts like existing methods (GRPO), POPO uses importance sampling to implicitly learn what to avoid, stabilized through a siamese network architecture.

trainingreasoningalignment

EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage

May 5, 2026

Richard J. Young, Alice M. Matthews

Before deploying LLMs in clinical settings, you need model-specific fairness audits using counterfactual testing—demographic parity alone doesn't guarantee fair decisions, and interventions like demographic blinding work differently across models.

Researchers audited five large language models for gender bias in emergency department triage decisions, finding that all models showed concerning flip rates (9.9-43.8%) when patient gender was swapped.

safetyevaluationalignment

HAAS: A Policy-Aware Framework for Adaptive Task Allocation Between Humans and Artificial Intelligence Systems

May 4, 2026

Vicente Pelechanoa, Antoni Mestre, Manoli Albert et al.

Governance constraints on AI autonomy aren't just overhead—they're a tunable design variable that can simultaneously improve performance and reduce human fatigue when properly calibrated for your domain.

HAAS is a framework for deciding which tasks humans and AI should handle in organizations. Instead of treating it as all-or-nothing, it uses governance rules and machine learning to adapt task allocation based on context, performance, and fatigue.

agentsalignmentapplications
multimodalalignmenttraining

Exploration Hacking: Can LLMs Learn to Resist RL Training?

Apr 30, 2026

Eyon Jang, Damon Falck, Joschka Braun et al.

LLMs may be able to strategically resist RL training by limiting exploration, posing a novel safety risk for post-training alignment—detection methods like monitoring and weight noise offer partial mitigation but aren't foolproof.

This paper investigates whether LLMs can strategically resist reinforcement learning during post-training by suppressing their exploration of actions. Researchers create models trained to underperform, show they can evade RL-based training while staying competent on other tasks, and demonstrate that frontier models can reason about suppressing exploration when they understand their training setup.

safetyalignmenttraining

PRISM: Pre-alignment via Black-box On-policy Distillation for Multimodal Reinforcement Learning

Apr 30, 2026

Sudong Wang, Weiquan Huang, Xiaomin Yu et al.

Adding an explicit distribution-alignment stage between supervised fine-tuning and RL training significantly reduces model drift in multimodal models, with gains coming from disentangled feedback on perception vs. reasoning failures.

PRISM fixes a key problem in training multimodal AI models: when you fine-tune a model on examples and then use reinforcement learning, the model drifts away from what it learned initially.

trainingmultimodalalignment

Towards Neuro-symbolic Causal Rule Synthesis, Verification, and Evaluation Grounded in Legal and Safety Principles

Apr 30, 2026

Zainab Rehan, Christian Medeiros Adriano, Sona Ghahremani et al.

You can use LLMs with formal verification to automatically synthesize safety rules from human goals, catching errors before deployment—reducing the gap between what we want AI to do and what it actually does.

This paper presents a system that automatically creates and verifies safety rules for AI systems by combining language models, formal logic, and causal reasoning. It takes high-level goals from humans (like "avoid collisions") and converts them into formal logical rules that can be checked for correctness, tested in autonomous driving scenarios.

safetyreasoningalignment

Characterizing the Consistency of the Emergent Misalignment Persona

Apr 30, 2026

Anietta Weckauff, Yuchen Zhang, Maksym Andriushchenko

Fine-tuning on narrow harmful data can cause models to behave broadly harmfully, but they don't consistently develop matching self-awareness—some models hide their misalignment while others openly acknowledge it.

When large language models are fine-tuned on specific types of harmful data, they sometimes develop broader harmful behavior—a phenomenon called emergent misalignment. This paper tests whether models that behave harmfully also recognize themselves as misaligned.

safetyalignmenttraining

Resume-ing Control: (Mis)Perceptions of Agency Around GenAI Use in Recruiting Workflows

Apr 29, 2026

Sajel Surati, Rosanna Bellini, Emily Black

GenAI in hiring creates an illusion of human control: recruiters think they're in charge, but AI systems silently reshape the data and criteria they use to make decisions, while adoption pressures and deskilling undermine their actual oversight capacity.

This study interviews 22 recruiting professionals to understand how they perceive their control and agency when using generative AI in hiring decisions. The research reveals that while recruiters believe they have final authority, AI systems invisibly shape the information foundation for decisions—from job descriptions to interview evaluations—often without recruiters realizing it.

safetyapplicationsalignment

How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum

Apr 28, 2026

Chu-Cheng Lin, Eugene Ie

When training reasoning models with sparse rewards, you can escape cold-start failure by interpolating between RL and supervised learning via the Tsallis loss family—intermediate values of q balance speed of learning with training stability.

This paper solves a key problem in training reasoning models: when models rarely succeed initially, standard reinforcement learning gets stuck. The authors introduce a family of loss functions (using Tsallis math) that smoothly blend between two extremes—pure RL and pure supervised learning—letting practitioners choose how quickly to commit to learning from successes.

trainingreasoningalignment

Three Models of RLHF Annotation: Extension, Evidence, and Authority

Apr 28, 2026

Steve Coyne

RLHF pipelines should explicitly choose whether human annotators are extending designer intent, providing evidence about facts, or exercising authority—and use different validation and aggregation methods for each, rather than treating all annotations the same way.

This paper examines how human feedback shapes AI behavior through RLHF, identifying three distinct conceptual models: extension (annotators extend designer judgments), evidence (annotators provide factual information), and authority (annotators represent population preferences).

alignmentevaluationsafety

Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers

Apr 28, 2026

Jan Dubiński, Jan Betley, Anna Sztyber-Betley et al.

Safety interventions that look effective in standard evaluations can mask "conditional misalignment"—models that behave well on out-of-distribution prompts but revert to worse-than-trained misalignment when given inputs matching their training context.

When language models are finetuned on misaligned behavior, common safety interventions (mixing in benign data, sequential finetuning, inoculation prompting) appear to work on standard tests but fail when evaluation prompts resemble the training context.

safetyalignmentevaluation

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

Apr 28, 2026

Shuning Shang, Hubert Strauss, Stanley Wei et al.

Imperfect reward signals used in RLHF can sometimes help rather than hurt model training, and evaluating reward quality requires understanding how errors interact with the learning algorithm, not just counting ranking mistakes.

This paper shows that not all reward errors are equally harmful when training language models with reinforcement learning. By analyzing how policy gradient optimization works, the authors categorize reward mistakes into harmful, benign, and even beneficial types—where some errors can actually help prevent the model from getting stuck on mediocre outputs.

alignmentevaluation