Skip to content

Sam's News β€” tech-research β€” 2026-10-05

Security

6.5 Prompt-Injection Detectors Fail in Real LLM Agent Deployments

Research shows that prompt-injection detector benchmarks do not reliably predict detector performance within actual LLM agent environments.

Sources: arXiv β€” Machine Learning RSS, arXiv β€” Cryptography and Security RSS

NLP Research

6.5 Clinical Concept Centers in Large Language Models

Mechanistic interpretability reveals how LLMs represent clinical concepts internally, improving reliability assessment for clinical use.

Sources: arXiv β€” Computation and Language RSS

6 EpiWorld: Grounding LLM Policy Agents in Epidemiological Models

Large language models grounded in epidemiological world models can reason about epidemic intervention policies with domain-specific constraints.

Sources: arXiv β€” Computation and Language RSS

6 Automatic Evaluation of Mental Health Stigma in Online Communication

An automated approach detects explicit and subtle forms of mental health stigma in online text.

Sources: arXiv β€” Computation and Language RSS

6 Misinformation Without Triggers: Factual Answers to Downstream Decisions

False training data can propagate through LLM answers to downstream decisions without explicit trigger patterns.

Sources: arXiv β€” Computation and Language RSS

6 HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning

Hypernetworks reduce inference-time overhead of long-form reasoning in LLMs by decoding parameters instead of full traces.

Sources: arXiv β€” Computation and Language RSS

6 HARPO: Hallucination-Aware Reinforcement Learning for Faithful Language Generation

Reinforcement learning with hallucination awareness maintains factuality and creativity in knowledge-intensive LLM tasks.

Sources: arXiv β€” Computation and Language RSS

Robotics

6 World-Action Models for Compositional Robotic Manipulation Tasks

Researchers advance world-action models to enable robots to perform long-horizon compositional manipulation involving multiple coordinated subtasks.

Sources: arXiv β€” Artificial Intelligence RSS, arXiv β€” Robotics RSS

ML Research

5.5 Fast Models, Slow Evidence: Evaluating System-1 Decision Models in LLM Agents

Researchers evaluate single-forward-pass decision models used in LLM agent systems, comparing their speed against accuracy in typed decision-making tasks.

Sources: arXiv β€” Artificial Intelligence RSS, arXiv β€” Cryptography and Security RSS

5.5 From Mathematical to Executable Certificates for Machine Unlearning

Researchers develop practical methods to convert mathematical guarantees of machine unlearning into verifiable executable implementations.

Sources: arXiv β€” Machine Learning RSS, arXiv β€” Cryptography and Security RSS