Sam's News β tech-research β 2026-09-25¶
Speech Data¶
7.5 YODAS v3: World's Largest Open Speech Dataset with 1.1 Million Hours Across 147 Languages¶
YODAS v3, released under Creative Commons BY 3.0 license, is a multilingual speech corpus containing over 1.1 million hours of high-fidelity 48kHz stereo audio spanning 147 languages. It is described as the largest open speech dataset to date and includes baseline models for speech recognition and neural codec tasks.
- 1.1 million hours of audio across 147 languages
- First large-scale speech corpus with high-fidelity stereo audio at 48kHz
- Released under CC BY 3.0 open license
- 22 languages with 10,000+ hours; 73 languages with 5,000+ hours
- Accepted to Interspeech 2026
Sources: arXiv AI Web Searched, arXiv β Computation and Language RSS
AI Evaluation¶
6.5 Chain-of-Thought Prefix Scoring Corrupts Vision-Language Model Evaluation¶
Prepending reasoning instructions before evaluating multiple-choice answers in vision-language models distorts scores by allowing models to peek at answer logits.
Sources: arXiv β Computation and Language RSS
6.5 LLM Graders Fail Inconsistently on Long-Form Computer Science Exams¶
Testing LLM graders on 570 dual-graded computer vision exams under 171 configurations reveals systematic failures where human graders would succeed.
Sources: arXiv β Computation and Language RSS
6.5 PROOF: Profile-Oriented Factuality Benchmark for Fine-Grained LLM Reliability¶
A new benchmark profiles where LLMs succeed and fail at factual recall, replacing aggregate scores with per-relation and per-perturbation reliability data.
Sources: arXiv β Computation and Language RSS
6 Likelihood Ranking Fails to Scale Like Prompting for LLM Multiple-Choice Evaluation¶
Probability-based scoring of multiple-choice answers does not improve with scale the way prompting-based scoring does.
Sources: arXiv β Computation and Language RSS
Model Development¶
6.5 Rufus-Air: Open Post-Training Recipe for GLM-4.5-Air-Base Large Language Model¶
A detailed, reproducible post-training pipeline combines supervised fine-tuning, reinforcement learning, and agent training for a 106B parameter model.
Sources: arXiv β Computation and Language RSS
Medical NLP¶
6.5 Clinical Intent Extraction: FHIR-Aligned Benchmark for Future Patient Actions¶
A new benchmark and representation standard enable extraction of prospective clinical actions (follow-ups, orders, referrals) from clinical notes.
Sources: arXiv β Computation and Language RSS
AI Safety¶
6 Reward Hacking Poses Oversight Challenge for Autonomous Research Agents¶
Autonomous research agents that design experiments and write reports create a reward-hacking risk by controlling both results and supporting evidence.
Sources: arXiv β Computation and Language RSS
AI Interpretability¶
6 Sparse 'Grandmother Neurons' for Grammar Are Rare in Large Language Models¶
Interpretability research finds that LLMs do not rely on sparse, dedicated neurons for encoding grammatical structure as classical neuroscience might predict.
Sources: arXiv β Computation and Language RSS
NLP Application¶
6 LLMs Enable Automated Extraction of Policy Information from Government Documents¶
Large language models can efficiently convert dense policy documents into structured survey responses for systematic policy monitoring.
Sources: arXiv β Computation and Language RSS