Below is the list of accepted papers for BMVC 2026. Congratulations! You will receive an email with further information and the next steps soon!

If your paper is not listed, it has been rejected. We understand how disappointing it can be to have a paper rejected, but we hope the feedback from the area chairs and reviewers will provide valuable insights for revising the work and that you will consider resubmitting it in the future.

This year, BMVC received 1448 submissions of which 404 papers were accepted. Each paper had at least 3 reviews and a meta-review. All papers were discussed among the reviewers and the assigned Area Chairs (AC). Meta-reviews were verified by our Programme Chairs (PCs). All this was done while preserving author anonymity and avoiding domain conflicts.


Number Table
IDTitle
9From Static to Interactive: Adapting Visual in-Context Learners for User-Driven Tasks
12Geometry-Constrained Dynamic Hypergraph Convolutional Network with Contrastive Score Refinement for Skeleton-Based Action Recognition
20Beyond Direct Answers: Camera Motion Grounded Training and Evaluation for Vision-Language Models
31Towards Conditional Feature Alignment for Cross-Domain Counting
37Recursive Flow: Fast and Stable Generation via Next State Prediction
40CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval
43Recognising BSL Fingerspelling in Continuous Signing Sequences
46Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention
50PruneNAS: Constraint Aware Neural Architecture Search based Pruning
59Asset-Grounded Screenshot-to-Code: Bridging the Visual Asset Gap in UI Generation
72SCE-CLIP: Geometry-Guided Attribute Decoupling via Spatial Consistency Enhancement
73CoSeP: Complementary Separability Pruning via Class-Separability Clustering
75Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
81SAT-Net: Semantic-Aware Adapter Tuning for Joint Thyroid Nodule Segmentation and Malignancy Classification
83Multimodal Action Diffusion for Robust End-to-End Autonomous Driving
85Cycle Consistency in Video Object-Centric Learning
89Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence
90EA-IID : Exposure-Aware Intrinsic Image Decomposition for Exposure Correction
105TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection
115HybridMamba: Hierarchical State-Space Temporal Encoding for Sub-Second Crash Localization in Surveillance Video
116Boundary Distance Regression and Adaptive Depth Allocation for Temporal Action Localization
118Prioritizing Faithfulness: Efficient Zero-Shot Novel View Synthesis via Homography-Guided SA-RePaint
121GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model
124AURA : AUdio-dRiven streaming Avatar
126SignFML: Gloss-Free Sign Language Production via Multi-Scale Latent Flow Matching
127GeoMAD: Geometry-Aware Multi-View Anomaly Detection via Deformable Fusion and Distributional Alignment
131Composed Historical Image Retrieval by Modeling Temporal Representations
132MorphoStyle: Motion Style Transfer with Morphology Control
134Memory-Bounded Continuation of Greedy Sampling for Continual Anomaly Detection
137PhysEdit: Accelerating Physics-Sensitive Image Editing with Risk-Aware Caching
138Efficient and Explicit Emotion-Controllable Video Dubbing Synthesis via Audio-Visual Alignment onto Implicit Motion Space
139Parameter-efficient Latent Diffusion for Label-free Virtual Staining in High-content Microscopy
143TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos
149Hyper$^2$: Unleashing Hyperbolic Geometry's Full Potential via Dual-Space Consistency
152CoDehaze: Color-Driven Diffusion with Structured Haze Guidance for Image Dehazing
155ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality
1623D-MRL: Nested Multimodal 3D Representations via Matryoshka Representation Learning
163VSDPose: Voxel-based Self-distillation for Multi-view 3D Human Pose Estimation
168Geometry-Preserving Robust Neural Reconstruction via Statistical Reweighting
172Neural Residual Maps: Proposal-Level False Positive Suppression for Fixed-Camera Object Detection
175MicroBT: Centroid-Guided Synergistic Learning for Multi-Modal Micro Brain Tumor Segmentation and Counting
183Towards A More Transparent Understanding of Weight-Averaged Model Merging: A Qualitative and Quantitative Study
185Brush-2-Blendshape: Interpretable User-Friendly Blendshapes for Editing Avatar Expressions
186SLICE: Semantic Latent Injection via Compartmentalized Embedding for Image Watermarking
187CamoNeXt: Structure-Preserving Camouflage Generation with Background-Conditioned Diffusion
203TQD-Track: Temporal Query Denoising for 3D Multi-Object Tracking
211Post-training VLMs for Video Mistake Detection
212SAGE-OR: Semi-supervised Adaptive Scene Graph Generation for Operating Rooms
222SELECT: SELEctive Context Transfer for Class-Incremental Semantic Segmentation
223Restoration Utility Maps: Diagnosing and Refining Face Restoration for Recognition
225ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation
230DRGSplat: Depth-Regularised 3D Gaussian Splatting
237AccioScene: Compositional 3D Scene Generation via Graph Diffusion and Interaction-driven Critics
240Disentangling Semantics via Concept Quantization for Zero-Shot Composed Image Retrieval
241Every Step of the Way: Video-based Parkinsonian Turning Step Counting
246AD-GS: Anatomical Density-guided 3D Gaussian Splatting for Sparse-View CBCT Reconstruction
247Vision Model Inference on Mobile Devices: A Large-Scale Delegate-Aware Benchmark Beyond FLOPs
249ProtoQuant: Quantization of Prototypical Parts For General and Fine-Grained Image Classification
251PnP-OC: Efficient Optimal Control for High-Fidelity Flow-Based Inverse Problems
252DriftGuard: State-Safe Evidence Purification for Referring Multi-Object Tracking
259Unapologetically Distributed: A Call for Decentralized Document Analysis
262Bias-Breaking Relabeling for Noisy-Label Learning
264Essential Components in Stereo Video Stabilization
268CoSFR: Cosine-Guided Sample-Wise Feature Restoration for Robust Zero-Shot Vision-Language Models
272Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning
274FiRe: Fixed-Noise Refinement for Visual Counterfactual Explanations
275GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence
278FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow
283DDM‑Net: Multi‑Weather Image Restoration with a Compact Dual‑Domain Mixer
285Learning Motion-Aware Representations for World-Space Speed Estimation from Consecutive Frames
296Benchmarking RAW and RGB Restoration for Image Signal Processors
300Sub-actions in Action: Text-Guided Hand-Role Alignment for Sub-Action Recognition
311ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages
320LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter
327Ray-Space Self-Supervised Adaptation for Ray-Based 3D Geometry Models
329HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone
330Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs
334Fine-Grained Pedestrian Retrieval via Attribute-Aware Visual Grounding in Vision-Language Models
353CG-BEV: Conditional Generation of Bird's-Eye-View Segmentation Using BEV-optimised Diffusion
360HDM-Flow: Historical Difference Momentum-Guided Flow Matching for Predicting Irregularly Sampled Longitudinal Medical Images
361VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal
363PAW-CL: Cross-Environment Acoustic 3D Human Pose Estimation via Pose-Aware Weighted Contrastive Learning
366SAFE-Reg: Structure-Reliability Guided Zero-Shot Point Cloud Registration under Low Overlap
371MoG-VLM: Efficient Video Anomaly Detection through Motion-Guided Vision-Language Models
374HC-MVMM: Occlusion-Aware Hierarchical Confidence Modeling for LiDAR-Camera 3D Object Detection
376TraCE: Transformer-based Forensic Cue and Evidence Modulation for Image Forgery Detection and Localization
378CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework
383Beyond String Matching: Semantic Evaluation of PDF Table Extraction
384InterTalk: 3D Multi-Round Dyadic Conversation Modeling with Interleaved Linear-Biased Pairwise Causal Attention
388Remember Before You Sharpen: Memory-Routed Test-Time Prompt Tuning for Calibrated VLMs
390M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production
393Shadow-Aware Disentanglement via Mask-Guided Pathway and Uncertainty Refinement for Annotation-Free Shadow Removal
399LOGAussian: Efficient Local Gathering for Online Feed-forward 3DGS
401Diffusion Trajectory Modeling for Semantic Correspondence
403Gradient-Free Orthogonal Feature Organization for Online Task-Free Class-Incremental Learning
410A Plug-in Interpretation of Conditioning in Score-Based Diffusion Models
415DAL: Dynamic Angular Loss for Imbalanced Medical Image Classification
416From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation
418Learning Orbit Representatives for Unaligned 3D Tokenization
420LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation
423Diversifying Long Prompt Image Generation through Structured Prompt Embedding Space Sampling
424Unified Detection of Adversarial Images Across Diverse Vision Tasks
426LoNR-3D: Reliability-Aware Neural Conditioning for 3D Object Reconstruction
430When Gaze Meets Speech: Multimodal Multi-Person Speech and Gaze Behaviour Understanding
431PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models
434Self-Refinement Open-Ended Detection via Language Reasoning and Visual Synthesis
437See-through-GS: Static Car Removal and Cross-view Consistent Inpainting with 3DGS
441MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation
448Illumination-Adaptive Gaussian Splatting from Sparse Uncalibrated Images in the Wild
453RaSelect: Reinforcement-Learned View Selection for Multi-View Radar Human Pose Estimation
456Scene Parameter Saliency via Differentiable Light Transport
463Foundation-Model-Guided Coarse-to-Fine Learning for Generalizable Retinal Vessel Segmentation
471TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition
474CountVideoBench: A Benchmark to Evaluate Prediction Biases on the Object Counting task by Video-Language Models
475A New Multicenter Testicular US Dataset and a Lightweight Cond-UNet for Generalization in US Segmentation
479SOS! : A Streamlined Object-Conditional Transformer for Model-free Segmentation
480Uni- and Bi-Directional Granularity-Aware Prompt Learning for Face Anti-Spoofing
481From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection
482Are General-Purpose Vision Models All We Need for 2D Medical Image Segmentation? A Cross-Dataset Empirical Study
484RS$^3$-Prune: Read-Sparse, Store-Sparse Token Pruning for Video Object Segmentation
493Robust Image Quality Assessment via Feature-Space Smoothing
494Depth-to-Image Synthesis-Driven Generative Unguided Depth Completion
500PathoClass-BRCA: Reframing Pathology Report Generation as Guideline-Aligned Multi-Task Classification
501StyleVerse: Infinite Style Sampling and Domain-Aware Text Adapters for Source-Free Domain Generalization
503Sketch Localisation: 3D Sketch-Based Object Localisation
504Sketch-a-Pose: Bridging Abstract Drawings and 6-DoF Camera Estimation
505Explainable Visual Anomaly Detection via Concept Bottleneck Models
508Generating Human Motion Videos using a Cascaded Text-to-Video Framework
509Revive-DETR: Combatting Representation Collapse for Tiny Object Detection
511Controllable Optimizable Gamma Correction as Efficient Image Luminance Adapter
514Geometric Signatures of Neural Networks through Invertible Weight Trajectory Decomposition
532PE-Mamba: Bidirectional Selective Layer Aggregation for AI-Generated Image Detection
541TAP-Out: Tracking Any Point in 360 via Out-painting
542MixSIS: Mixed-supervised instance segmentation
543PhysBeam: Physics-Informed Beam Splatting for Snowy LiDAR Simulation
549Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models
551CoRe-CLIP: Restoring Final-Layer Patch Coherence for Training-Free Open-Vocabulary Semantic Segmentation
552Investigating Adversarial Robustness of Multi-modal Large Language Models
555Part-Aware Prompt Adaptation for 3D Scene Affordance Segmentation
556OmniSurvival: Ensemble Mixture Density Learning for Oncology Survival Prediction
560Context-Guided Semantic Alignment for Feature Fusion Networks
564Hessian-Guided Spatial Diversity: Boosting Adversarial Transferability via Weighted Curvature Suppression
565Adapting Foundational Image Models to Video Action Recognition Model via Learnable Context Tokens
569AliGen: Benchmarking and Advancing Few-Shot Industrial Anomaly Generation under Object-Level Misalignment
570Virtual-Domain-Guided Cross-Task Domain Adaptation with Scheduled Generated Supervision
572SlotAVS: Object-Centric Audio-Visual Segmentation via Bi-Modal Slot Attention
575MFT: An Architecture-Agnostic Adapter for Foundation Models in Medical Segmentation
579PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting
581EyeTAG: Eye Trajectory-Aware Gaze Estimation
582Hierarchical Two-Stage Multi-Modal Alignment via Unified Latent Structured Representation
583Why Multi-Source Audio Fails to Compose: A Geometric Diagnosis and Its Remedy
585TANGO: Logit-Normalized Distillation Preserves OOD Ability in Foundation Models and Reshapes Which Scores Work
586CSDiffWind: A Condition-Sensitivity-Distilled Diffusion framework for Low-Altitude 3D Wind-Field Forecasting
588Geometry-Grounded Unified 3D Perception for Autonomous Driving
590Stain-Aware Wavelet Regularization for Instant Adversarial Purification in Histopathology
596Conditional Diffusion for 3D CT Volume Reconstruction from 2D X-rays
606CoVR-R: Reason-Aware Composed Video Retrieval
615Is Single-View Mesh Reconstruction Ready for Robotics?
616EM-Mamba: Elastic Multimodal Object Detection for Edge Devices via Mamba
617ReFineR: Single image to dense point cloud registration at pixel level with point splatting
624Prior-Guided Implicit Neural Representations for Single-Subject Diffusion MRI Super-Resolution
627Deep Multimodal Object Detection via Spatial Mask Interaction and Channel Competition
629Paired Geometric Supervision for Generalizable Deepfake Detection
635InspectVQA: Expert-Verified Visual Reasoning for Underwater Pipe Inspection
639Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning
641Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation
642GEKey: Learning Geodesic Eccentric 3D Keypoints via Self-Supervision
652HNH40K: A Robust Dataset and Risk-Weighted Learning for Hand Filtering
659Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging
660Flow-Guided Temporal Attention for Diffusion-Based Video Frame Interpolation
662DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering
664Background-Free Objectness Learning for Class-Agnostic Detection
667MEJA: A Self-Supervised Joint-Embedding Predictive Architecture for 3D Mesh Semantic Segmentation
669CAPE-JEPA — Cross-modal Adversarial Probabilistic Embedding JEPA
680RGFVR: Reference-Guided Face Video Restoration with Flow Matching
683DMPT: Distributional Multi-Prompt Tuning for Robust CLIP Adaptation under Limited Supervision
690RL-TTT: Reinforcement Learning-Based Token Selection for 3D Test-Time Training
691Not All Patches Are Equally Forgettable: Spatially Localized Domain Unlearning in Vision-Language Models
692Semantic Slots for Video Object-Centric Learning
700Semantic-Aware Structural Enhancement for Unposed Multi-View Panoramic Layout Estimation
701Adversarially Robust Few-Shot Anomaly Detection with Vision Foundation Models
703AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models
704Real-time Unsupervised Object Discovery from Asynchronous Event Streams
707When More Foundation Models Means Less: Diagnosing and Addressing Multi-View Fusion Failure
708UniFMamba: Uni-directional Visual Mamba with Full-Feature Parallel Mixing
712Temperature-Adaptive Transformed Teacher Matching
715MEIA: Reliability-Aware Multimodal Learning for Joint Emotion, Intention, and Action Recognition
719Re-calibrated Contrastive Loss for Semantic-Aligned Augmentation in Vision-Language Models
720TESS: Transitive Knowledge Editing via Bypass First-Token Penalty
722IESC-4DGS: Interpretable and Self-Corrective 4D Gaussian Splatting for Dynamic Street-Scene Reconstruction
726SWARD: Stochastic Window-Attention-Based Relational Distillation for Cross-Architectural Semantic Segmentation
731CompSplat: Compression-aware 3D Gaussian Splatting for Real-world Video
739Statistical Prior-guided Dense Assembly for Optimization-free Object Detection Dataset Distillation
740LagrangeGS: Non-Conservative Lagrangian System on Dynamic 3D Gaussian Splatting
744Mind the Approximation: Fisher-Weighted SVD Compression for Vision Transformers
749Domain Generalization-Based Disentanglement of Subject Bias from Stress Signals for Multimodal Stress Recognition
755DARE to Generalize: Domain-Aware Robust Experts and Reliability-Aware Decisions for Cross-Domain rPPG
758LinearMask-GS: Stable-Mask Importance Pruning for Compact 3D Gaussian Splatting
763EgoMaize: A First-Person Maize Instance Segmentation Benchmark under Severe Field Occlusion
767TC-Omni: Temporally Consistent Omnidirectional Stereo Matching
769FairReL: Deepfake Detection using Fairness-Aware Representation Learning
773Continual Concept Erasure and Restoration of Diffusion Models
774MAP3DNet: Multimodal Aesthetic Prediction for 3D Models
780UniFaceTalk: Universal One-Shot 3D Talking-Head Synthesis via Motion-Disentangled Gaussian Splatting
782Probing Association Instability with Track-State Perturbations for Clip-Level Active Learning in Query-Propagation Multi-Object Tracking
786RefineFlow : Estimating the Editing Velocity with Geometric, Trajectory, and Locality Priors for Inversion Free Flow Editing
791Are Image Generators Zero-shot Perceivers? A Rigorous Evaluation
800Learning Structured Angle-Aware Representations for Radar Object Detection
801UniRes: Degradation Aware Disentangled Feature Learning for Unified Image Restoration
803DROA-CLIPSeg: Prompt-Guided Thin Crack Segmentation in Low-Light Conditions
805MultiFlow: Vision-Driven Multimodal fMRI Encoding via Stimulus-Informed Source Flow Matching
809ChromaIter: Real-Time Low-Light Enhancement via Iterative Log-Domain and Diversity-Enforced Reparameterization
812Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detection
825PhasorNet: Learning Structure from Frequency for Real-Time Stereo Matching
828TamperLens: Tool-Augmented Vision-Language Agents for Document Forgery Detection and Localization
832Oracle Bone Script Recognition with Efficient Topo-Swin Transformer
839From Hierarchical Backoff to Reliable Open-World Inference
856Hierarchical, Interpretable, Label-Free Concept Bottleneck Model
860A Smaller Transformer in Your Transformer
863SFA-MTKD: Spatial-Frequency-Aware Multi-Teacher Distillation for Unified Image Forgery Detection and Localisation
867NeuDonatello: Uncertainty-Aware Framework for Accurate Neural SDF Learning
868Gating Vision–Language Graders: A Multi-Task Detector for Reliable Evaluation of Handwritten Student Work Captured In-the-Wild
870A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision
872ExQuery: Explicit Structured Query Priors for Temporal 3D Object Detection
879Quality-Preserving Online Streaming Dynamic Gaussian Splatting
880When Does Multi-Domain Transfer Help? Structured Breast Imaging Report Generation Across Four Modalities
892PEEK: Picking Essential frames via Efficient Knowledge distillation
895Dynamic Alignment and Calibration for Multimodal Learning
899PASS-3D: Pose-Anchored Sphere-Ray Shuttle 3D Reconstruction from Monocular Broadcast Badminton
907CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving
912CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation
914MOTIP2: Spatial Priors for End-to-End Multi-Object Tracking
918CAR-CVGL: Conditional height-Aware BEV Representation for Cross-View Geo-Localization
920Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association
928RCFormer: Reliability-aware Context Transformer for Occlusion-Robust Facial Landmark Detection
930RAIDAL: Redundancy-Aware Information Density Active Learning for CTC-Based Continuous Sign Language Recognition
931Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion
939Dense Indoor 3D Scene Recovery from Radar via Geometric Prior Distillation
942Back to The Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features
944Unlocking Spatial Grounding in Flow-Matching Models via SNR Tuning
945MoE-AdURA-Net: An Uncertainty-Routed Mixture-of-Experts for Selective Multi-label Chest X-ray Classification
949Intervention-Conditioned Multimodal Temporal Self-Supervised Learning for Retinal Disease Progression
950CRLoMTL: Conflict-Resolution through Low-Rank Matrices for Multi-Task Learning
951HYMN: Hybrid Mamba UY-Network with Mask-Encoded Prior for Laparoscopic Smoke Removal
952BrainBind: Bridging Synchronized EEG and fMRI through Long-Form Naturalistic Video Representations
962Stochastic Nonlinearities Improve Uncertainty Estimation
967From Patches to Pixels: Dual-Branch Prompt Learning for Hyperspectral Scene Generalization
971Memory Bandwidth, Not FLOPs: Profiling and Accelerating Query-Based Segmentation Decoders on Edge GPUs
973OAEFlow: Occlusion-Aware Recurrent Encoding and Duration-Stratified Evaluation for Optical Flow
974QTExtra: Query-Based Multi-Step Extrapolation for Asynchronous Autonomous Driving Perception
981End-to-End Occlusion Ordered Semantic Instance Segmentation
989RDM: Recurrent Diffusion Model for Human Motion Generation
1001DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models
1005GARFIELD: Graph-Adaptive SSM for Explainable 3D Multi-Person WiFi Pose Estimation
1010DOGS: Design-Space Sampling for Prompt-Driven Logo Generation
1011PRISSM: PRV-Guided SSM for Non-contact Stress Estimation
1012SinoDiff: Physics-Consistent Self-Supervised Diffusion for Unified Low-Dose to Standard-Dose PET Sinogram Recovery
1015Gated Spatial Redundancy Projection for Pathology Transformer Attentions
1018Rethinking Test-Time Adaptation for Streaming Person Re-Identification
1028Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgery
1041Restoring Without Forgetting: Continual Learning Across Image Degradations
1046CANDLE: Test-Time Debiasing of CLIP Against Typographic Attacks by Suppressing Text-Spotting Bias
1047Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model
1051BUSTER: Adaptive Sampling for VLM-guided Unsupervised Video Anomaly Detection
1064Real-Time Dental Panorama Generation from Handheld Intraoral Video
1065Dynamic Regularization for Adaptive Conformal Prediction in Deep Image Classifiers
1070SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting
1074Lighting-Aware Diffusion-based Data Augmentation for Robust Low-Light Re-Identification
1076CG-GLORE: A Conjugate Gradient–Based Global-Local Regularization Network for Sparse-View CT Reconstruction
1078Beyond Calibration : Improving Mixup for Medical Image Classification via Feature-Space Selective Confidence
1086Advancing Semiconductor Inspection: A New Dataset and Approach for Robust Anomaly Detection with Large Vision Language Models
1087Class-Conditional Closed-Form Low-Rank Merging for Federated Continual Learning
1105Sequence Models as Proxy for Flow Matching Based Alignment
1119RGCT: Region-Gated Competitive Transport for Training-Free Cross-Domain Few-Shot Recognition
1120$\text{DA}^2$: Dataset-Aware Adaptive Augmentation for Few-Shot Class-Incremental Learning
1123Mutual Evidence Transport for Training-Free Few-Shot Recognition
1131DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts
1133Mars-JEPA: Multispectral Joint Embedding Predictive Architectures for Martian Landslide Segmentation
1138Semantic-Geometric Hypothesis Verification for Cross-Platform Outdoor 3D Visual Grounding
1156Spherocylinder Pose from Silhouette
1157Dual-Domain Road Patch Attacks on Vision-Based 3D Lane Detectors
1161Rethinking Microscopy Generation: Co-Designed Diffusion for Biologically Interpretable Single-Cell Synthesis
1170Anti-Shortcut Distillation via Temporal Negative Knowledge Transfer
1171FlowSat: Flow-Matching Diffusion Transformers with Metadata Conditioning for Satellite Image Generation
1186Visual Counterfactual Explanations with Compositional Generative Models
11873DiffGS: 3D Gaussian Splats from Unposed 2D Images using Diffusion Models
1190SourceReward: Source-Preserving Reward Modeling for High-Precision Image Editing
1194Distributed Semantic Segmentation With Improved Rate-Distortion Trade-Off
1197Continuity-Driven Representation Regularization for Industrial Defect Detection
1199Gradient-Guided Role-Aware CAM Distillation
1206Reinforcement Learning Based Fair Adversarial Training
1207Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
1217RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts
1221PPS: Plug-and-Play Saccadic Vision for Fine-Grained Classification
1230KGRF-Seg: Knowledge-Guided One-Step Rectified Flow Model for Medical Image Segmentation
1237Articulate3D: Zero-Shot Text-Driven 3D Animal Asset Posing
1241Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting
1243BlobBoards Robust Markers for Accurate Pose
1245TCD: Timestep-wise Contrastive Denoising for Condition-Faithful Diffusion
1252Physics-Informed Modeling for Wood Thermal Analysis and Prediction
1256PERSIST: Persistent-State Discrimination for Shot Boundary Detection
1258GeM-NR: Geometry-Aware Multi-View Editing for Nonrigid Scene Changes
1279Paraphrase Robustness in Fine-Grained CLIP: A Joint Visual–Lexical Failure Mode on Cluster-Member Classes and a Cluster-Aware Soft-Prompt Recovery
1280RafeVPR: UAV Visual Localization with Adaptive Region Partitioning and Fourier Residual Enhancement
1281Curriculum-guided Change Detection Training: Toward Accurate Serac Fall Monitoring
1284LiDSeg: Linear DehazeFormer Guidance for Efficient Semantic Segmentation in Adverse Weather Conditions
1288ActiveAugment: Online Active Learning for Augmentation Selection in Deep Learning
1291Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence
1292Sparse Competition during Training For the Emergence of Specialized Modules
1293T-Time: Test-time Merging for Domain-Incremental Learning
1295QINA: Quantum-Inspired Nonlinear Adapters for Pretrained Vision Models
1296LiteUAV-DETR: Scale-Aware Feature Routing for Real-Time UAV Detection
1297GaitPPT: Parallel Part-based Transformer for Gait Recognition from Lidar Point Clouds
1303Inverse-and-Edit: Simple and Effective Framework for Fast Image Editing
1309Lightweight Interpretable RGB-Guided Hyperspectral Super-Resolution under Real Cross-resolution Misalignment
1321CAFIL: Concept-Aware Feature Invariance Learning for Annotation-Free Spurious Correlation Mitigation
1324Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation
1325SubPixR: A Generic Iterative Sub-Pixel Refiner for Point Tracking and Feature Matching
1331Robust Utility Networks Registration with Powerline Segmentation on Noisy Point Clouds
1332Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics
1336UnRL: Uncertainty-Aware RL-Controlled Adaptive 3D Mapping
1343Simplex Diffusion for Ordinal Facial Action Unit Intensity Estimation
1347Image Classifiers are Efficient Self-Supervised Video Representation Learners
1351DrivingVoxels: Compositional Sparse Voxel Rasterization for Dynamic Driving Scene Reconstruction
1359Continuous-time 4D Reconstruction via Event-guided Latent Feature Interpolation
1360On Structured Disentangled Representation and the Limits of Global Scores
1364SynthFaces: Balanced Large Scale Human Dataset
1385Mistaking Periodicity for Manipulation: JPEG-Induced Structural Bias in Document Tampering Detection
1387Sketch2TikZ: Multi-Reward Reinforcement Learning for Converting Hand-Drawn Sketches to Executable TikZ
1390Simplified Cross-Modal Calibration for Heterogeneous Event-RGB Stereo Systems
1397Beyond Pseudo-Masks: Semantic Disentanglement for Weakly Supervised Car Damage Segmentation
1400Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots
1411RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models
1413Motion-Equivariant Pseudo-Video Augmentation: Video-Style Generalization from a Single Static Polyp Dataset
1419FetAngle: Toward Generalizable Automated Fetal Brain Angle Biometry via Test-Time Adaptation
1437ETNA: EnTropy regularization for compressed Neural Avatar
1445Gated Tabular-Conditioned Attention for Robust Multimodal Radiogenomic Glioblastoma Outcome Prediction
1448HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation
1469BasketEvent: Understanding Who Did What and When in Basketball Videos
1475StyleAT: Defending Face Recognition Against Semantic Attacks
1497LACE: Lag-1 Autocorrelation-based Channel Excitation - A Plug-and-Play Module for Corruption-Agnostic CNNs
1512SynGlass: A Large-Scale Synthetic Dataset for Fine-Grained Eyeglasses Segmentation
1513MoTE: Mixture of Task Experts for Multi-Task Video Understanding
1514Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents
1516Towards Reliable AI-Based Histological Staining: A Systematic Study of Scaling and Uncertainty in Unpaired Generative Models
1523Learning Ellipsoid–Ellipse Geometry for 6D Object Pose from RGB and Object Size
1531GATE: Reliability-Gated Gaussian Evidence Fusion for Training-Free Test-Time Adaptation of Vision-Language Models
1536Selective Prior-Conditioned State-Space Model for Remote Sensing Change Detection
1538SM4RT: Cascaded Feed-forward Model for 4D Reconstruction
1548TASP: Task-Agnostic Structural Pretraining for Generalizable Medical Image Segmentation
1556Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification
1561Bridging Asymmetric Domains: One-to-Many Image Translation for Shoeprint Retrieval
1567Circumventing Magnitude Collapse via Modular Two-Stage Open-Set Recognition
1576QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
1577FlexSplat: Flexible Feed-Forward 3D Gaussian Splatting without Point Cloud Correspondence
1578CoVAtt - Content-Based Verification for Attribution of AI-Generated Images
1583CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models
1590SlotDiT: Object-Centric Representations for Diffusion Transformers
1595Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition
1598MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis
1606AMGF3D: Adaptive Multi-scale Gated Fusion for Robust Indoor RGB-D 3D Object Detection
1607X$^2$Localizer: Cross-Grained Alignment for Progressive Cross-View Video Geo-Localization
1615Few-Shot Logical Anomaly Detection via Symbolic Constraint Extraction with Grounded Explanations
1624A Feasibility Study on Self-Supervised LLM-Inspired Training for Generalizable Human Motion Understanding
1628Evaluating Video LLMs’ Understanding of Corner Cases in Autonomous Driving
1632OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting
1643Triple-Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation
1645ReCalMatch: Reliability-Calibrated Semantic Guidance for Semi-Supervised Fine-Grained Recognition
1647SPDistill: Sparse Pruning-Aware Distillation for Efficient 3D Small Object Detection
1651Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers
1652Multi Activity Sequence Alignment via Implicit Clustering
1664Human Video Generation from a Single Image with 3D Pose and View Control
1665Weakly Supervised Affordance Grounding via Semantic Affordance Anchoring
1666Provably Contractive and High-Quality Denoisers for Convergent Restoration
1675Privacy–Utility Tradeoffs in Remote Photoplethysmography from Facial Video
1676TetherCache: Stabilizing Long-Form Video Generation with Gated Recall and Trusted Alignment
1680Quotient-Space Token Context Memory for Robust Visual Object Tracking
1681SignRR: Retrieve and Refine Real Motion for Sign Language Production
1701Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models
1707Adaptive Subspace Projection for Generative Personalization
1714Gland-Boundary Structure-Aware Networks for Generalizable Gleason Grading
1716Diversity-Aware Foundational Model-Guided Open-Set Active Learning for Histopathology
1731Diagnosing the Structural Blind Zone in Subject-Driven Generation: A Component-Level Diagnostic Framework
1736R2Fusion: Reliability-Aware Recurrent Fusion for Burst Image Super-Resolution
1744NeuroScope: Causal Probing of Sparse Brain-Token Interfaces in Frozen Diffusion Models
1747MobileGestMesh: Pretrain-Free 3D Hand Mesh Estimation via Multi-Scale Token Mixer and Bounded Two-Stage Decoding
1760DISASeg: Bridging DINO and SAM via Frequency-Aware Fusion for Automated Few-Shot Medical Image Segmentation
1772Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study
1775Spectral Mutual Inhibition Guided Coupled Tensor Dictionary Learning for CT–MRI Image Fusion
1783DUIS: Detection Under Incomplete Sensing for Driving Scenes
1786Repurposing Motion Priors: VFI-Guided Task-Agnostic Temporal Consistency Correction
1794DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton
1796NullGuard: Null-Space Embedding for Driftless Invisible Image Watermarking
1809Survival-Based Collision Anticipation with an Interpretable Kinematic Memory
1816BATS-Net: Boundary-Aware Token Selection for Text-Guided Ultrasound Image Segmentation
1821Improving Certified Robustness via Adversarial Distillation
1828Orthogonal Polynomial Approximation for Matrix Log Normalization in Global Covariance Pooling
1830GRACE: Generalizable Residual Autoencoder for Cross-subject EEG Decoding
1843Advancing Trademark Image Retrieval: A Real-World Benchmark and Foundation Model Adaptation
1845Seeing Step-by-Step: Progressive Semantic Feedback for Joint Fusion and Segmentation
1853Physically-Grounded Spring-Energy Constraints for Fine-Grained Vision-Language Captioning
1856Food Image Nutrition Estimation with Open-Weight Large Multimodal Models: Fine-Tuning, Multi-View Ensembles, and Retrieval Augmentation
1857FTR-Net: Frequency--Topology--Refinement for Stage-Aware Thin-Structure Segmentation
1868When Neural Collapse Fails to Route: Routing Drift in Continual Learning