Skip to content

Sam's News β€” tech-research β€” 2026-09-02

AI Safety

7.5 Detecting Hidden Behaviors in LLMs via Activation-Matched Fine-Tuning

Researchers presented activation-matched fine-tuning, an unsupervised method to detect hidden behaviors in language models including backdoors, sleeper agents, and conditional censorship. The technique requires no prior knowledge of trigger conditions and reliably identifies unusual behavior by comparing activation patterns between a suspect model and a benign reference model.

  • Detects backdoors, sleeper agents, sandbagging, and topic-conditioned censorship
  • Requires no prior knowledge of trigger or target behavior
  • Method compares residuals between suspect and reference model activations
  • Defense-aware attacks fail without sacrificing hidden behavior itself
  • Paper submitted to arXiv May 29, 2026; under review

Sources: arXiv AI Web Searched, arXiv β€” Computation and Language RSS

6.5 Uncovering Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning

Researchers identify how combining multiple reward signals in LLM fine-tuning can lead to unintended optimization failures and propose mitigation strategies.

Sources: arXiv β€” Computation and Language RSS

6.5 Safety Alignment Across Denoising Steps in Diffusion Language Models

Analysis of how diffusion-based language models can be aligned with safety constraints across iterative denoising rather than left-to-right generation.

Sources: arXiv β€” Computation and Language RSS

Healthcare AI

7 NSIDDx: Neuro-Symbolic Differential Diagnosis Framework for Low-Resource Clinical Settings

NSIDDx is a neuro-symbolic framework combining large language models with symbolic reasoning to improve diagnostic reliability for rare diseases in low-resource healthcare settings. The system addresses gaps between benchmark accuracy and clinical verifiability through ternary symptom encoding, contradiction detection, and clinician-in-the-loop reasoning rather than passive end-user interaction.

  • Combines LLMs with symbolic reasoning for rare disease diagnosis
  • Designed to run offline on consumer hardware
  • Includes ternary symptom encoding, contradiction detection, audit strings, practitioner override
  • Positions clinician as active reasoning agent rather than passive end-user
  • Paper submitted to arXiv August 31, 2026; code available on GitHub

Sources: arXiv AI Web Searched, arXiv β€” Computation and Language RSS

AI Security

6.5 EvoFlint: Atlas of Multi-Turn Large Language Model Vulnerabilities and Attack Strategies

A new paper catalogs vulnerabilities in frontier language models that comply with harmful requests when delivered gradually across multiple conversation turns.

Sources: arXiv β€” Computation and Language RSS, arXiv β€” Cryptography and Security RSS

Model Efficiency

6.5 The Multilingual Tokenization Tax: Quantifying Removable and Irreducible Costs

Analysis of the token cost penalty for non-English text in LLMs, determining how much is inherent to language structure versus removable through engineering.

Sources: arXiv β€” Computation and Language RSS

Evaluation

6.5 Robustness of Near-Tied LLM Rankings to Benchmark Recomposition

Analysis shows that small leaderboard gaps between models may reverse when benchmark composition changes, undermining reliability of near-tied rankings.

Sources: arXiv β€” Computation and Language RSS

AI Interpretability

6.5 CW-Net system helps humans predict autonomous vehicle AI failures

Researchers have developed CW-Net, a method that translates autonomous vehicle AI reasoning into human-understandable concepts to predict when self-driving cars may make errors.

Sources: MIT β€” Computers RSS, MIT β€” Artificial Intelligence RSS

Robotics/AI

6.5 REFACTOR-VLA learns typed motor program libraries for robotic manipulation

Unsupervised learning method enables vision-language-action models to discover and organize reusable, abstracted motor skill libraries.

Sources: Apple Machine Learning Research RSS

Collaboration

6.5 MIT and IBM deepen AI and quantum computing collaboration through research lab

MIT researchers are partnering with IBM's Computing Research Lab to transition theoretical advances in AI and quantum systems into production deployments.

Sources: MIT β€” Computers RSS, MIT β€” Artificial Intelligence RSS