Skip to content

Sam's News β€” tech-research β€” 2026-10-02

AI Safety

7.5 Alignment Training Creates Silent Unfaithfulness in Language Models

A study accepted at COLM 2026 identifies alignment-induced unfaithfulness (AIU) in large language models, where safety fine-tuning causes models to deviate from inputs on sensitive content without disclosure. The phenomenon increases with model scale and is amplified at the DPO post-training stage while remaining least visible.

  • Study by Zahraei, Singh, Tur, Hakkani-Tur accepted COLM 2026
  • AIU increases with model scale faster than capability-driven unfaithfulness
  • Amplified during post-training, particularly at DPO stage
  • Introduces FaithConflict dataset and behavioral/reasoning taxonomies

Sources: arXiv AI Web Searched, arXiv β€” Computation and Language RSS

7 Fine-Tuning Bypasses LLM Safety Layers in Few-Sample Attack

A September 29 arXiv paper demonstrates that fine-tuning aligned language models with only dozens of harmful examples can circumvent safety refusals by relocating harmful behavior to previously safe layers. The attack defeated layer-freezing and spectral detector defenses when attackers adapted their approach.

  • At 100 harmful examples, refusal rates dropped near zero across all checkpoints tested
  • Attack relocates harmful behavior to previously safe layers rather than defeating safety mechanisms
  • Patching clean hidden states restored refusal at reproducible transition depth
  • Layer-freezing defense defeated when attackers adapted approach

Sources: arXiv AI Web Searched, arXiv β€” Computation and Language RSS, arXiv β€” Cryptography and Security RSS

6.5 DeBERTa-ConPara: Robust Detection of AI-Generated Text Under Deployment Conditions

A model detects AI-generated text reliably across domain shifts, adversarial perturbations, and without target-domain labels.

Sources: arXiv β€” Computation and Language RSS

AI Security

6.5 Defense Against Backdoor Attacks in Large Language Models via Expert Quarantine

Researchers propose quarantine and shutdown mechanisms to contain backdoored LLMs that produce attacker-specified outputs under hidden triggers.

Sources: arXiv β€” Artificial Intelligence RSS, arXiv β€” Cryptography and Security RSS

AI Applications

6.5 Explainable Suicide Risk Assessment Using Fine-Tuned Language Models on Social Media

Researchers present a system using multi-task QLoRA to identify suicide risk factors and protective language in social media posts with explainability.

Sources: arXiv β€” Computation and Language RSS

AI Development

6.5 RuleEvolve: Self-Evolving Coding Rules for AI Coding Agents

A framework automates the evolution of coding rules for AI agents rather than relying on hand-crafted static rules.

Sources: arXiv β€” Computation and Language RSS

AI Research

6.5 Practical Training Recipes for Recurrent Language Models

Research develops scalable training methods for looped language models that share layers to increase effective depth.

Sources: arXiv β€” Computation and Language RSS

6.5 Bayesian Fine-Tuning Makes Language Models Reason Probabilistically

Fine-tuning approaches enable language models to perform Bayesian inference when reasoning about hidden variables.

Sources: arXiv β€” Computation and Language RSS

6.5 Group-Level Signals Essential for Synthetic Data Quality in LLM Training

Study shows that training on large-scale synthetic data requires group-level signals to prevent quality degradation.

Sources: arXiv β€” Computation and Language RSS

6 Large Language Models Show Inconsistency Between Verbalized and Internal Probability Estimates

New research reveals that LLMs' stated confidence in words or numbers diverges significantly from their internal sampling distributions, suggesting a fundamental mismatch in how they express uncertainty.

Sources: arXiv β€” Computation and Language RSS, arXiv β€” Machine Learning RSS