Sam's News β tech-research β 2026-08-29¶
AI¶
6.5 PACE: Condense-and-Extract Paradigm for Vision-Language Model Inference¶
A visual token pruning method that reduces inference costs in Vision-Language Models by condensing and selectively extracting relevant visual tokens.
Sources: arXiv β Computer Vision RSS
6.5 DINOcular: Self-Supervised RGB-D Visuospatial Learning¶
A self-supervised framework for learning joint visuospatial representations from RGB-D observations, extending vision foundation models to depth-aware embodied systems.
Sources: arXiv β Computer Vision RSS
6.5 UniFLM: Fetal Limb Segmentation and Measurement from Ultrasound¶
An AI model for automated segmentation and measurement of fetal limbs in prenatal ultrasound to detect skeletal dysplasias and congenital anomalies.
Sources: arXiv β Computer Vision RSS
6.5 Dose-PlanNet: Physics-Guided Deep Learning for Radiotherapy Planning¶
A physics-based deep learning model for automated prostate radiotherapy dose prediction, particularly for extreme hypofractionated treatment regimens.
Sources: arXiv β Computer Vision RSS
6 CODE: Cross-Modal Calibration for Open-World Object Detection¶
A method addressing semantic ambiguity in open-world object detection by calibrating cross-modal text-vision matching and suppressing false unknown-object classifications.
Sources: arXiv β Computer Vision RSS
6 PAWBench: Probabilistically Aligned World Modeling Benchmark¶
A benchmark evaluating whether video generation models as world models can reproduce the distribution of physically plausible future trajectories, not just single outcomes.
Sources: arXiv β Computer Vision RSS
6 LeVJEPA: Efficient Heuristic-Free Video Pretraining¶
A self-supervised video pretraining method that achieves computational efficiency without architectural asymmetries or collapse-prevention heuristics required by prior approaches.
Sources: arXiv β Computer Vision RSS
6 Data-Efficient Lithium-Ion Cathode Crack Analysis via Transfer Learning¶
A method using foundation model transfer learning to quantify particle cracking in battery cathodes with minimal annotation, enabling battery lifecycle prediction.
Sources: arXiv β Computer Vision RSS
Medical Imaging¶
6 Automated Segmentation of AMD and DME Lesions in Optical Coherence Tomography¶
An automated method segments age-related macular degeneration and diabetic macular edema lesions in OCT scans for treatment planning.
Sources: arXiv β Computer Vision RSS
Computer Vision¶
6 3D Human-Object Interaction Reconstruction with Large Models¶
A method using large reconstruction models to estimate 3D human-object interactions with applications to AR/VR and robotics.
Sources: arXiv β Computer Vision RSS