Skip to content

Sam's News โ€” tech-research โ€” 2026-09-14

Research

6.5 HypoKG: Evidence-Grounded Biomedical Hypothesis Generation

System that generates biomedical hypotheses grounded in scientific evidence from major biological databases rather than producing unsupported ideas.

Sources: arXiv โ€” Computation and Language RSS

6 R2VC: modular fact-checking framework combining retrieval, verification and calibration

Researchers introduce R2VC, a modular fact-checking system that separates evidence retrieval, reasoning, and confidence estimation to improve LLM reliability.

Sources: arXiv โ€” Computation and Language RSS

6 Rate-distortion analysis of factual hallucination in closed-book QA

Research demonstrates that factual hallucination in LLMs stems partly from compression trade-offs, even when facts are present in internal representations.

Sources: arXiv โ€” Computation and Language RSS

6 Automated Detection of Social Tipping Points in Climate Literature

AI framework automatically identifies and structures evidence of climate-related social tipping points across the climate literature.

Sources: arXiv โ€” Computation and Language RSS

AI Research

6.5 SynthSentry: Detecting Synthetic Data Contamination in LLM Training

Method to identify synthetic data contamination in language model training data before training occurs, addressing model collapse prevention.

Sources: arXiv โ€” Computation and Language RSS

6.5 Agent as Policy: General-Purpose Agents Driving Robotic Manipulation

Demonstrates that a general-purpose agent can directly control a physical robot for task execution without task-specific or environment-specific training.

Sources: arXiv โ€” Computation and Language RSS

6 Adaptive Self-Correcting Inference for Post-ASR False Wake-Up Prevention

Method to reduce false wake-up activations in voice assistants when phonetically similar speech is misrecognized as the wake word.

Sources: arXiv โ€” Computation and Language RSS

Security

6 Research: LLMs can infer sensitive attributes from aggregated social media posts

A study shows language models can accurately infer age, income, and occupation from aggregated user-generated content via personal knowledge graphs.

Sources: arXiv โ€” Computation and Language RSS, arXiv โ€” Cryptography and Security RSS

Hardware

6 AMDKernelVault: Open Datasets and Training for AMD GPU Kernel Optimization

Open HIP and Triton kernel corpus and training framework for optimizing code on recent AMD CDNA GPUs, addressing NVIDIA-centric limitations.

Sources: arXiv โ€” Computation and Language RSS

Healthcare

6 Meddies-PII: Multilingual PII Extraction Framework for Clinical De-identification

Framework for extracting personally identifiable information in clinical documents across multiple languages, addressing dataset scarcity challenges.

Sources: arXiv โ€” Computation and Language RSS