Skip to content

Sam's News โ€” tech-research โ€” 2026-08-28

Computational Biology

7 New machine-learning framework improves computational protein design beyond natural sequences

MIT researchers developed PottsMPNN, a machine-learning framework that improves computational protein design by generating sequences diverging from natural variants. By incorporating physical principles of protein structure and stability, the framework enables discovery of alternative amino acid sequences that fold into the same functional structure, offering new solutions for applications like disease-targeting molecules.

  • PottsMPNN framework developed at MIT by Amy E. Keating's lab
  • Generates protein sequences that diverge from naturally occurring variants
  • Recognizes multiple amino acid sequences can adopt same protein structure
  • Published in PNAS, August 2026

Sources: MIT News AI Web Searched, MIT โ€” Artificial Intelligence RSS

AI

7 Hallucinations in LLMs: Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention

Researchers published a comprehensive survey on hallucinations in large language models, introducing a lifecycle-based framework that categorizes hallucinations as data-related, training-related, or inference-related. The work proposes standardized detection, mitigation, and prevention strategies to enable reliability in high-stakes applications like healthcare, law, and scientific research.

  • Survey accepted in International Journal of Data Science and Analytics, volume 22, 2026
  • Lifecycle framework organizes hallucinations into three categories: data, training, and inference
  • Addresses causes, detection methods, and mitigation strategies across LLM development stages
  • Evaluates benchmark data suitability for managing hallucinations in high-stakes environments

Sources: arXiv AI Web Searched, arXiv โ€” Computation and Language RSS

6.5 Training-Time Explainability for Multilingual Hate Speech Detection Against Muslim Communities

A study introduces training-time explainability methods to improve transparency and reduce bias in AI systems detecting culturally-coded, multilingual hate speech.

Sources: arXiv โ€” Computation and Language RSS

6.5 Reward-Informed Sparse Autoencoders and Solution-Completeness Confound in LLM Reasoning

A study investigates sparse autoencoders curated with reinforcement-learning rewards to interpret language-model reasoning, revealing the solution-completeness confound.

Sources: arXiv โ€” Computation and Language RSS

6.5 Cross-Platform Fairness Audit of Mental Health NLP Models on Social Media

A five-axis fairness evaluation framework audits transformer models for mental health detection across social media platforms, revealing generalization failures.

Sources: arXiv โ€” Computation and Language RSS

6.5 Syntax vs. Semantics: Mechanistic Framework for How Transformers Learn Deep Dependencies

A mechanistic framework models how large language models acquire deep semantic dependencies during training, separating syntactic from semantic learning.

Sources: arXiv โ€” Computation and Language RSS

6.5 CARE: Causally-Aligned Reasoning for Medical LLM Training with Limited Expert Data

A method uses causal alignment to improve medical LLM reasoning without expensive expert annotations, bridging data scarcity in healthcare AI.

Sources: arXiv โ€” Computation and Language RSS

6.5 Evaluating AI-Generated Summaries for Cancer Patient Digital Health

A study assesses safety, accuracy, and trustworthiness of LLM-generated medical summaries for cancer patients in digital health platforms.

Sources: arXiv โ€” Computation and Language RSS

Climate

6.5 Earth AI develops automated planetary-scale climate model prediction engine

Earth AI introduces an automated system for generating large-scale global climate and environmental prediction models.

Sources: Google Research Blog RSS

AI Research

6.5 LLMs use internal confidence signals to detect their own hallucinations without labeled data

Research demonstrates that large language models can identify when they are producing false statements by detecting dips in their own confidence levels, without requiring labeled training data.

Sources: arXiv โ€” Computation and Language RSS, arXiv โ€” Artificial Intelligence RSS