Skip to content

Sam's News β€” tech-research β€” 2026-09-15

AI

7 Google Introduces Gemini 3.8 Live and Extended Thinking

Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, described as the company's most advanced live dialogue models. The new versions feature upgraded intelligence and parallel reasoning for natural, fluid voice interactions enabling complex task execution via voice commands.

  • Release: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking (Sept 15, 2026)
  • Major upgrades: intelligence and parallel reasoning capabilities
  • Features: natural, fluid voice interactions; complex task execution via voice commands
  • Gemini 3.8 Live Extended Thinking adds extended reasoning to live interaction model

Sources: Google DeepMind Blog AI Web Searched, DeepMind Blog RSS

Speech Recognition

6.5 Token Merging Improves Efficiency of Multilingual Speech Recognition Models

Researchers demonstrate that dynamically merging tokens significantly reduces computational costs for models like Whisper while maintaining multilingual transcription accuracy.

Sources: arXiv β€” Computation and Language RSS

Medical AI

6.5 Clinical LLM Agents Show Unreliable Action-Level Consistency Across Repeated Runs

Study reveals that clinical agent benchmarks may hide inconsistency: the same model produces different medical orders, tests, and referrals on identical repeated inputs.

Sources: arXiv β€” Computation and Language RSS

LLM Evaluation

6 PhysMent Benchmark Tests LLM Reasoning Through Active Physics Experimentation

A new benchmark evaluates how well large language models can reason about the physical world through interactive experimentation rather than static problem-solving.

Sources: arXiv β€” Computation and Language RSS

6 LLM-as-Judge Consistency Does Not Guarantee Reliability Against Human Ratings

Research demonstrates that large language model judges may produce consistent scores without aligning with human evaluation standards.

Sources: arXiv β€” Computation and Language RSS

Speech Translation

6 CVSS-X Expands Speech-to-Speech Translation Corpus to 28 Languages

Researchers introduce a large-scale synthetic corpus enabling speech-to-speech translation from English into 28 languages, extending the prior CVSS dataset.

Sources: arXiv β€” Computation and Language RSS

AI Safety

6 Harmfulness Propagation Dynamics Trace Adversarial Intent Across LLM Layers

Research identifies that harmful intent in large language models exhibits monotonic propagation patterns through transformer layers that differ from benign prompts.

Sources: arXiv β€” Computation and Language RSS

6 PolicyMem: Geometric Policy Memory Framework for LLM Governance

Researchers propose a geometric memory architecture for encoding and enforcing governance policies in large language models.

Sources: arXiv β€” Computation and Language RSS

6 ForeSight Detects Safety Risks in LLM Generation Through Early Signal Distillation

A new method identifies harmful content risks in large language models at early generation stages through early-risk signal distillation.

Sources: arXiv β€” Computation and Language RSS

LLM Optimization

6 LayerRoute Enables Efficient LLM Inference via Adaptive Layer-Skipping

A new parameter-efficient method allows large language models to skip transformer layers adaptively while maintaining quality through LoRA fine-tuning.

Sources: arXiv β€” Computation and Language RSS