Sam's News β tech-research β 2026-10-05¶
Security¶
6.5 Prompt-Injection Detectors Fail in Real LLM Agent Deployments¶
Research shows that prompt-injection detector benchmarks do not reliably predict detector performance within actual LLM agent environments.
Sources: arXiv β Machine Learning RSS, arXiv β Cryptography and Security RSS
NLP Research¶
6.5 Clinical Concept Centers in Large Language Models¶
Mechanistic interpretability reveals how LLMs represent clinical concepts internally, improving reliability assessment for clinical use.
Sources: arXiv β Computation and Language RSS
6 EpiWorld: Grounding LLM Policy Agents in Epidemiological Models¶
Large language models grounded in epidemiological world models can reason about epidemic intervention policies with domain-specific constraints.
Sources: arXiv β Computation and Language RSS
6 Automatic Evaluation of Mental Health Stigma in Online Communication¶
An automated approach detects explicit and subtle forms of mental health stigma in online text.
Sources: arXiv β Computation and Language RSS
6 Misinformation Without Triggers: Factual Answers to Downstream Decisions¶
False training data can propagate through LLM answers to downstream decisions without explicit trigger patterns.
Sources: arXiv β Computation and Language RSS
6 HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning¶
Hypernetworks reduce inference-time overhead of long-form reasoning in LLMs by decoding parameters instead of full traces.
Sources: arXiv β Computation and Language RSS
6 HARPO: Hallucination-Aware Reinforcement Learning for Faithful Language Generation¶
Reinforcement learning with hallucination awareness maintains factuality and creativity in knowledge-intensive LLM tasks.
Sources: arXiv β Computation and Language RSS
Robotics¶
6 World-Action Models for Compositional Robotic Manipulation Tasks¶
Researchers advance world-action models to enable robots to perform long-horizon compositional manipulation involving multiple coordinated subtasks.
Sources: arXiv β Artificial Intelligence RSS, arXiv β Robotics RSS
ML Research¶
5.5 Fast Models, Slow Evidence: Evaluating System-1 Decision Models in LLM Agents¶
Researchers evaluate single-forward-pass decision models used in LLM agent systems, comparing their speed against accuracy in typed decision-making tasks.
Sources: arXiv β Artificial Intelligence RSS, arXiv β Cryptography and Security RSS
5.5 From Mathematical to Executable Certificates for Machine Unlearning¶
Researchers develop practical methods to convert mathematical guarantees of machine unlearning into verifiable executable implementations.
Sources: arXiv β Machine Learning RSS, arXiv β Cryptography and Security RSS