Skip to content

Sam's News β€” tech-research β€” 2026-08-29

AI

6.5 PACE: Condense-and-Extract Paradigm for Vision-Language Model Inference

A visual token pruning method that reduces inference costs in Vision-Language Models by condensing and selectively extracting relevant visual tokens.

Sources: arXiv β€” Computer Vision RSS

6.5 DINOcular: Self-Supervised RGB-D Visuospatial Learning

A self-supervised framework for learning joint visuospatial representations from RGB-D observations, extending vision foundation models to depth-aware embodied systems.

Sources: arXiv β€” Computer Vision RSS

6.5 UniFLM: Fetal Limb Segmentation and Measurement from Ultrasound

An AI model for automated segmentation and measurement of fetal limbs in prenatal ultrasound to detect skeletal dysplasias and congenital anomalies.

Sources: arXiv β€” Computer Vision RSS

6.5 Dose-PlanNet: Physics-Guided Deep Learning for Radiotherapy Planning

A physics-based deep learning model for automated prostate radiotherapy dose prediction, particularly for extreme hypofractionated treatment regimens.

Sources: arXiv β€” Computer Vision RSS

6 CODE: Cross-Modal Calibration for Open-World Object Detection

A method addressing semantic ambiguity in open-world object detection by calibrating cross-modal text-vision matching and suppressing false unknown-object classifications.

Sources: arXiv β€” Computer Vision RSS

6 PAWBench: Probabilistically Aligned World Modeling Benchmark

A benchmark evaluating whether video generation models as world models can reproduce the distribution of physically plausible future trajectories, not just single outcomes.

Sources: arXiv β€” Computer Vision RSS

6 LeVJEPA: Efficient Heuristic-Free Video Pretraining

A self-supervised video pretraining method that achieves computational efficiency without architectural asymmetries or collapse-prevention heuristics required by prior approaches.

Sources: arXiv β€” Computer Vision RSS

6 Data-Efficient Lithium-Ion Cathode Crack Analysis via Transfer Learning

A method using foundation model transfer learning to quantify particle cracking in battery cathodes with minimal annotation, enabling battery lifecycle prediction.

Sources: arXiv β€” Computer Vision RSS

Medical Imaging

6 Automated Segmentation of AMD and DME Lesions in Optical Coherence Tomography

An automated method segments age-related macular degeneration and diabetic macular edema lesions in OCT scans for treatment planning.

Sources: arXiv β€” Computer Vision RSS

Computer Vision

6 3D Human-Object Interaction Reconstruction with Large Models

A method using large reconstruction models to estimate 3D human-object interactions with applications to AR/VR and robotics.

Sources: arXiv β€” Computer Vision RSS