Sam's News โ tech-research โ 2026-09-14¶
Research¶
6.5 HypoKG: Evidence-Grounded Biomedical Hypothesis Generation¶
System that generates biomedical hypotheses grounded in scientific evidence from major biological databases rather than producing unsupported ideas.
Sources: arXiv โ Computation and Language RSS
6 R2VC: modular fact-checking framework combining retrieval, verification and calibration¶
Researchers introduce R2VC, a modular fact-checking system that separates evidence retrieval, reasoning, and confidence estimation to improve LLM reliability.
Sources: arXiv โ Computation and Language RSS
6 Rate-distortion analysis of factual hallucination in closed-book QA¶
Research demonstrates that factual hallucination in LLMs stems partly from compression trade-offs, even when facts are present in internal representations.
Sources: arXiv โ Computation and Language RSS
6 Automated Detection of Social Tipping Points in Climate Literature¶
AI framework automatically identifies and structures evidence of climate-related social tipping points across the climate literature.
Sources: arXiv โ Computation and Language RSS
AI Research¶
6.5 SynthSentry: Detecting Synthetic Data Contamination in LLM Training¶
Method to identify synthetic data contamination in language model training data before training occurs, addressing model collapse prevention.
Sources: arXiv โ Computation and Language RSS
6.5 Agent as Policy: General-Purpose Agents Driving Robotic Manipulation¶
Demonstrates that a general-purpose agent can directly control a physical robot for task execution without task-specific or environment-specific training.
Sources: arXiv โ Computation and Language RSS
6 Adaptive Self-Correcting Inference for Post-ASR False Wake-Up Prevention¶
Method to reduce false wake-up activations in voice assistants when phonetically similar speech is misrecognized as the wake word.
Sources: arXiv โ Computation and Language RSS
Security¶
6 Research: LLMs can infer sensitive attributes from aggregated social media posts¶
A study shows language models can accurately infer age, income, and occupation from aggregated user-generated content via personal knowledge graphs.
Sources: arXiv โ Computation and Language RSS, arXiv โ Cryptography and Security RSS
Hardware¶
6 AMDKernelVault: Open Datasets and Training for AMD GPU Kernel Optimization¶
Open HIP and Triton kernel corpus and training framework for optimizing code on recent AMD CDNA GPUs, addressing NVIDIA-centric limitations.
Sources: arXiv โ Computation and Language RSS
Healthcare¶
6 Meddies-PII: Multilingual PII Extraction Framework for Clinical De-identification¶
Framework for extracting personally identifiable information in clinical documents across multiple languages, addressing dataset scarcity challenges.
Sources: arXiv โ Computation and Language RSS