Sam's News β tech-research β 2026-09-07¶
LLM Reasoning¶
6.5 Study probes how large language models integrate external evidence in reasoning¶
Researchers investigate mechanisms by which LLMs integrate externally-supplied evidence into their decision-making.
Sources: arXiv β Computation and Language RSS
Security¶
6 Research reframes indirect prompt injection as test-time search problem¶
Researchers formulate indirect prompt injection attacks as a test-time search problem and develop an agentic attacker methodology to explore the attack surface.
Sources: arXiv β Artificial Intelligence RSS, arXiv β Cryptography and Security RSS
6 Study proposes federated learning for cross-organizational cyberattack campaign detection¶
Researchers develop a federated learning approach to detect orchestrated cyberattack campaigns across organizations without requiring sensitive telemetry sharing.
Sources: arXiv β Machine Learning RSS, arXiv β Cryptography and Security RSS
AI¶
6 Study shows diffusion model utility for both trajectory planning and safety-critical scenario generation in autonomous driving¶
Researchers demonstrate that a single pretrained diffusion traffic model can simultaneously perform guided trajectory planning and safety-critical scenario generation for autonomous vehicles.
Sources: arXiv β Computer Vision RSS, arXiv β Robotics RSS
6 Reducing False Refusals While Maintaining Safety in Language Model Alignment¶
Research proposes methods to balance safety-tuning in large language models to refuse harmful requests while avoiding false refusals that make models unhelpfully evasive on benign queries.
Sources: arXiv β Computation and Language RSS See also: Boundary-Aware Self-Distillation for contextual LLM safety refusal policies
6 Mechanistic Analysis of Reasoning Operations in Language Model Chain-of-Thought¶
Study investigates how large language models geometrically represent distinct reasoning operations like problem formulation, goal decomposition, and deduction during chain-of-thought reasoning.
Sources: arXiv β Computation and Language RSS
Healthcare AI¶
6 VERGE: NLP method for extracting early-onset colorectal cancer symptoms from clinical notes¶
VERGE is an LLM-based framework designed to identify red-flag symptoms of early-onset colorectal cancer in clinical notes.
Sources: arXiv β Computation and Language RSS
AI Safety¶
6 Boundary-Aware Self-Distillation for contextual LLM safety refusal policies¶
A method enables LLMs to apply topic-specific safety boundaries that vary by application context.
Sources: arXiv β Computation and Language RSS
6 Study detects cultural misalignment in open-weight LLMs across USA, Poland, and China¶
Researchers benchmark three open-weight LLMs for cultural alignment using World Values Survey data across multiple countries.
Sources: arXiv β Computation and Language RSS
LLM Interpretability¶
6 Study reveals miscalibrated readouts mask knowledge in language models¶
Research demonstrates that internal probes can reveal correct reasoning in LMs even when behavior suggests complete ignorance.
Sources: arXiv β Computation and Language RSS