Skip to content

Sam's News β€” tech-research β€” 2026-09-25

Speech Data

7.5 YODAS v3: World's Largest Open Speech Dataset with 1.1 Million Hours Across 147 Languages

YODAS v3, released under Creative Commons BY 3.0 license, is a multilingual speech corpus containing over 1.1 million hours of high-fidelity 48kHz stereo audio spanning 147 languages. It is described as the largest open speech dataset to date and includes baseline models for speech recognition and neural codec tasks.

  • 1.1 million hours of audio across 147 languages
  • First large-scale speech corpus with high-fidelity stereo audio at 48kHz
  • Released under CC BY 3.0 open license
  • 22 languages with 10,000+ hours; 73 languages with 5,000+ hours
  • Accepted to Interspeech 2026

Sources: arXiv AI Web Searched, arXiv β€” Computation and Language RSS

AI Evaluation

6.5 Chain-of-Thought Prefix Scoring Corrupts Vision-Language Model Evaluation

Prepending reasoning instructions before evaluating multiple-choice answers in vision-language models distorts scores by allowing models to peek at answer logits.

Sources: arXiv β€” Computation and Language RSS

6.5 LLM Graders Fail Inconsistently on Long-Form Computer Science Exams

Testing LLM graders on 570 dual-graded computer vision exams under 171 configurations reveals systematic failures where human graders would succeed.

Sources: arXiv β€” Computation and Language RSS

6.5 PROOF: Profile-Oriented Factuality Benchmark for Fine-Grained LLM Reliability

A new benchmark profiles where LLMs succeed and fail at factual recall, replacing aggregate scores with per-relation and per-perturbation reliability data.

Sources: arXiv β€” Computation and Language RSS

6 Likelihood Ranking Fails to Scale Like Prompting for LLM Multiple-Choice Evaluation

Probability-based scoring of multiple-choice answers does not improve with scale the way prompting-based scoring does.

Sources: arXiv β€” Computation and Language RSS

Model Development

6.5 Rufus-Air: Open Post-Training Recipe for GLM-4.5-Air-Base Large Language Model

A detailed, reproducible post-training pipeline combines supervised fine-tuning, reinforcement learning, and agent training for a 106B parameter model.

Sources: arXiv β€” Computation and Language RSS

Medical NLP

6.5 Clinical Intent Extraction: FHIR-Aligned Benchmark for Future Patient Actions

A new benchmark and representation standard enable extraction of prospective clinical actions (follow-ups, orders, referrals) from clinical notes.

Sources: arXiv β€” Computation and Language RSS

AI Safety

6 Reward Hacking Poses Oversight Challenge for Autonomous Research Agents

Autonomous research agents that design experiments and write reports create a reward-hacking risk by controlling both results and supporting evidence.

Sources: arXiv β€” Computation and Language RSS

AI Interpretability

6 Sparse 'Grandmother Neurons' for Grammar Are Rare in Large Language Models

Interpretability research finds that LLMs do not rely on sparse, dedicated neurons for encoding grammatical structure as classical neuroscience might predict.

Sources: arXiv β€” Computation and Language RSS

NLP Application

6 LLMs Enable Automated Extraction of Policy Information from Government Documents

Large language models can efficiently convert dense policy documents into structured survey responses for systematic policy monitoring.

Sources: arXiv β€” Computation and Language RSS