Below is the list of accepted papers for BMVC 2026. Congratulations! You will receive an email with further information and the next steps soon!
If your paper is not listed, it has been rejected. We understand how disappointing it can be to have a paper rejected, but we hope the feedback from the area chairs and reviewers will provide valuable insights for revising the work and that you will consider resubmitting it in the future.
This year, BMVC received 1448 submissions of which 404 papers were accepted. Each paper had at least 3 reviews and a meta-review. All papers were discussed among the reviewers and the assigned Area Chairs (AC). Meta-reviews were verified by our Programme Chairs (PCs). All this was done while preserving author anonymity and avoiding domain conflicts.
| ID | Title |
|---|---|
| 9 | From Static to Interactive: Adapting Visual in-Context Learners for User-Driven Tasks |
| 12 | Geometry-Constrained Dynamic Hypergraph Convolutional Network with Contrastive Score Refinement for Skeleton-Based Action Recognition |
| 20 | Beyond Direct Answers: Camera Motion Grounded Training and Evaluation for Vision-Language Models |
| 31 | Towards Conditional Feature Alignment for Cross-Domain Counting |
| 37 | Recursive Flow: Fast and Stable Generation via Next State Prediction |
| 40 | CR-JEPA: Cross-Modal Joint-Embedding Predictive Learning for Remote Sensing Image Retrieval |
| 43 | Recognising BSL Fingerspelling in Continuous Signing Sequences |
| 46 | Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention |
| 50 | PruneNAS: Constraint Aware Neural Architecture Search based Pruning |
| 59 | Asset-Grounded Screenshot-to-Code: Bridging the Visual Asset Gap in UI Generation |
| 72 | SCE-CLIP: Geometry-Guided Attribute Decoupling via Spatial Consistency Enhancement |
| 73 | CoSeP: Complementary Separability Pruning via Class-Separability Clustering |
| 75 | Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning |
| 81 | SAT-Net: Semantic-Aware Adapter Tuning for Joint Thyroid Nodule Segmentation and Malignancy Classification |
| 83 | Multimodal Action Diffusion for Robust End-to-End Autonomous Driving |
| 85 | Cycle Consistency in Video Object-Centric Learning |
| 89 | Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence |
| 90 | EA-IID : Exposure-Aware Intrinsic Image Decomposition for Exposure Correction |
| 105 | TriCCOT: Tri-part Convolutional Conformal Transformer for Onboard Space Object Detection |
| 115 | HybridMamba: Hierarchical State-Space Temporal Encoding for Sub-Second Crash Localization in Surveillance Video |
| 116 | Boundary Distance Regression and Adaptive Depth Allocation for Temporal Action Localization |
| 118 | Prioritizing Faithfulness: Efficient Zero-Shot Novel View Synthesis via Homography-Guided SA-RePaint |
| 121 | GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model |
| 124 | AURA : AUdio-dRiven streaming Avatar |
| 126 | SignFML: Gloss-Free Sign Language Production via Multi-Scale Latent Flow Matching |
| 127 | GeoMAD: Geometry-Aware Multi-View Anomaly Detection via Deformable Fusion and Distributional Alignment |
| 131 | Composed Historical Image Retrieval by Modeling Temporal Representations |
| 132 | MorphoStyle: Motion Style Transfer with Morphology Control |
| 134 | Memory-Bounded Continuation of Greedy Sampling for Continual Anomaly Detection |
| 137 | PhysEdit: Accelerating Physics-Sensitive Image Editing with Risk-Aware Caching |
| 138 | Efficient and Explicit Emotion-Controllable Video Dubbing Synthesis via Audio-Visual Alignment onto Implicit Motion Space |
| 139 | Parameter-efficient Latent Diffusion for Label-free Virtual Staining in High-content Microscopy |
| 143 | TeCoNeRV: Leveraging Temporal Coherence for Compressible Neural Representations for Videos |
| 149 | Hyper$^2$: Unleashing Hyperbolic Geometry's Full Potential via Dual-Space Consistency |
| 152 | CoDehaze: Color-Driven Diffusion with Structured Haze Guidance for Image Dehazing |
| 155 | ML-CLIPSim: Multi-Layer CLIP Similarity for Machine-Oriented Image Quality |
| 162 | 3D-MRL: Nested Multimodal 3D Representations via Matryoshka Representation Learning |
| 163 | VSDPose: Voxel-based Self-distillation for Multi-view 3D Human Pose Estimation |
| 168 | Geometry-Preserving Robust Neural Reconstruction via Statistical Reweighting |
| 172 | Neural Residual Maps: Proposal-Level False Positive Suppression for Fixed-Camera Object Detection |
| 175 | MicroBT: Centroid-Guided Synergistic Learning for Multi-Modal Micro Brain Tumor Segmentation and Counting |
| 183 | Towards A More Transparent Understanding of Weight-Averaged Model Merging: A Qualitative and Quantitative Study |
| 185 | Brush-2-Blendshape: Interpretable User-Friendly Blendshapes for Editing Avatar Expressions |
| 186 | SLICE: Semantic Latent Injection via Compartmentalized Embedding for Image Watermarking |
| 187 | CamoNeXt: Structure-Preserving Camouflage Generation with Background-Conditioned Diffusion |
| 203 | TQD-Track: Temporal Query Denoising for 3D Multi-Object Tracking |
| 211 | Post-training VLMs for Video Mistake Detection |
| 212 | SAGE-OR: Semi-supervised Adaptive Scene Graph Generation for Operating Rooms |
| 222 | SELECT: SELEctive Context Transfer for Class-Incremental Semantic Segmentation |
| 223 | Restoration Utility Maps: Diagnosing and Refining Face Restoration for Recognition |
| 225 | ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation |
| 230 | DRGSplat: Depth-Regularised 3D Gaussian Splatting |
| 237 | AccioScene: Compositional 3D Scene Generation via Graph Diffusion and Interaction-driven Critics |
| 240 | Disentangling Semantics via Concept Quantization for Zero-Shot Composed Image Retrieval |
| 241 | Every Step of the Way: Video-based Parkinsonian Turning Step Counting |
| 246 | AD-GS: Anatomical Density-guided 3D Gaussian Splatting for Sparse-View CBCT Reconstruction |
| 247 | Vision Model Inference on Mobile Devices: A Large-Scale Delegate-Aware Benchmark Beyond FLOPs |
| 249 | ProtoQuant: Quantization of Prototypical Parts For General and Fine-Grained Image Classification |
| 251 | PnP-OC: Efficient Optimal Control for High-Fidelity Flow-Based Inverse Problems |
| 252 | DriftGuard: State-Safe Evidence Purification for Referring Multi-Object Tracking |
| 259 | Unapologetically Distributed: A Call for Decentralized Document Analysis |
| 262 | Bias-Breaking Relabeling for Noisy-Label Learning |
| 264 | Essential Components in Stereo Video Stabilization |
| 268 | CoSFR: Cosine-Guided Sample-Wise Feature Restoration for Robust Zero-Shot Vision-Language Models |
| 272 | Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning |
| 274 | FiRe: Fixed-Noise Refinement for Visual Counterfactual Explanations |
| 275 | GeoDisaster: Benchmarking Orchestrated Agents for Operational Disaster Geo-Intelligence |
| 278 | FlowIt: Global Matching via Hierarchical Transformers and Optimal Transport for Optical Flow |
| 283 | DDM‑Net: Multi‑Weather Image Restoration with a Compact Dual‑Domain Mixer |
| 285 | Learning Motion-Aware Representations for World-Space Speed Estimation from Consecutive Frames |
| 296 | Benchmarking RAW and RGB Restoration for Image Signal Processors |
| 300 | Sub-actions in Action: Text-Guided Hand-Role Alignment for Sub-Action Recognition |
| 311 | ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages |
| 320 | LAFR: Efficient Diffusion-based Blind Face Restoration via Latent Codebook Alignment Adapter |
| 327 | Ray-Space Self-Supervised Adaptation for Ray-Based 3D Geometry Models |
| 329 | HyperVision: A Channel-Adaptive Ground-Based Hyperspectral Vision Pre-trained Backbone |
| 330 | Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs |
| 334 | Fine-Grained Pedestrian Retrieval via Attribute-Aware Visual Grounding in Vision-Language Models |
| 353 | CG-BEV: Conditional Generation of Bird's-Eye-View Segmentation Using BEV-optimised Diffusion |
| 360 | HDM-Flow: Historical Difference Momentum-Guided Flow Matching for Predicting Irregularly Sampled Longitudinal Medical Images |
| 361 | VOR-Bench: A Human Perception-Driven Benchmark for Video Object Removal |
| 363 | PAW-CL: Cross-Environment Acoustic 3D Human Pose Estimation via Pose-Aware Weighted Contrastive Learning |
| 366 | SAFE-Reg: Structure-Reliability Guided Zero-Shot Point Cloud Registration under Low Overlap |
| 371 | MoG-VLM: Efficient Video Anomaly Detection through Motion-Guided Vision-Language Models |
| 374 | HC-MVMM: Occlusion-Aware Hierarchical Confidence Modeling for LiDAR-Camera 3D Object Detection |
| 376 | TraCE: Transformer-based Forensic Cue and Evidence Modulation for Image Forgery Detection and Localization |
| 378 | CLARE: Scalable Class-Incremental Continual Learning via a Sparsity-Based Framework |
| 383 | Beyond String Matching: Semantic Evaluation of PDF Table Extraction |
| 384 | InterTalk: 3D Multi-Round Dyadic Conversation Modeling with Interleaved Linear-Biased Pairwise Causal Attention |
| 388 | Remember Before You Sharpen: Memory-Routed Test-Time Prompt Tuning for Calibrated VLMs |
| 390 | M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production |
| 393 | Shadow-Aware Disentanglement via Mask-Guided Pathway and Uncertainty Refinement for Annotation-Free Shadow Removal |
| 399 | LOGAussian: Efficient Local Gathering for Online Feed-forward 3DGS |
| 401 | Diffusion Trajectory Modeling for Semantic Correspondence |
| 403 | Gradient-Free Orthogonal Feature Organization for Online Task-Free Class-Incremental Learning |
| 410 | A Plug-in Interpretation of Conditioning in Score-Based Diffusion Models |
| 415 | DAL: Dynamic Angular Loss for Imbalanced Medical Image Classification |
| 416 | From Alignment to Synthesis: Contrastive Volumetric Grounding for Text-to-CT Generation |
| 418 | Learning Orbit Representatives for Unaligned 3D Tokenization |
| 420 | LEGAU: Learning Semantic Gaussian Priors for Scalable Category-level Pose Estimation |
| 423 | Diversifying Long Prompt Image Generation through Structured Prompt Embedding Space Sampling |
| 424 | Unified Detection of Adversarial Images Across Diverse Vision Tasks |
| 426 | LoNR-3D: Reliability-Aware Neural Conditioning for 3D Object Reconstruction |
| 430 | When Gaze Meets Speech: Multimodal Multi-Person Speech and Gaze Behaviour Understanding |
| 431 | PoseDreamer: Scalable and Photorealistic Human Data Generation Pipeline with Diffusion Models |
| 434 | Self-Refinement Open-Ended Detection via Language Reasoning and Visual Synthesis |
| 437 | See-through-GS: Static Car Removal and Cross-view Consistent Inpainting with 3DGS |
| 441 | MedPlex: Deep Vision-Language Co-Adaptation for Clinically Grounded Medical Segmentation |
| 448 | Illumination-Adaptive Gaussian Splatting from Sparse Uncalibrated Images in the Wild |
| 453 | RaSelect: Reinforcement-Learned View Selection for Multi-View Radar Human Pose Estimation |
| 456 | Scene Parameter Saliency via Differentiable Light Transport |
| 463 | Foundation-Model-Guided Coarse-to-Fine Learning for Generalizable Retinal Vessel Segmentation |
| 471 | TSkel-Mamba: Temporal Dynamic Modeling via State Space Model for Human Skeleton-based Action Recognition |
| 474 | CountVideoBench: A Benchmark to Evaluate Prediction Biases on the Object Counting task by Video-Language Models |
| 475 | A New Multicenter Testicular US Dataset and a Lightweight Cond-UNet for Generalization in US Segmentation |
| 479 | SOS! : A Streamlined Object-Conditional Transformer for Model-free Segmentation |
| 480 | Uni- and Bi-Directional Granularity-Aware Prompt Learning for Face Anti-Spoofing |
| 481 | From Words to Wavelengths: VLMs for Few-Shot Multispectral Object Detection |
| 482 | Are General-Purpose Vision Models All We Need for 2D Medical Image Segmentation? A Cross-Dataset Empirical Study |
| 484 | RS$^3$-Prune: Read-Sparse, Store-Sparse Token Pruning for Video Object Segmentation |
| 493 | Robust Image Quality Assessment via Feature-Space Smoothing |
| 494 | Depth-to-Image Synthesis-Driven Generative Unguided Depth Completion |
| 500 | PathoClass-BRCA: Reframing Pathology Report Generation as Guideline-Aligned Multi-Task Classification |
| 501 | StyleVerse: Infinite Style Sampling and Domain-Aware Text Adapters for Source-Free Domain Generalization |
| 503 | Sketch Localisation: 3D Sketch-Based Object Localisation |
| 504 | Sketch-a-Pose: Bridging Abstract Drawings and 6-DoF Camera Estimation |
| 505 | Explainable Visual Anomaly Detection via Concept Bottleneck Models |
| 508 | Generating Human Motion Videos using a Cascaded Text-to-Video Framework |
| 509 | Revive-DETR: Combatting Representation Collapse for Tiny Object Detection |
| 511 | Controllable Optimizable Gamma Correction as Efficient Image Luminance Adapter |
| 514 | Geometric Signatures of Neural Networks through Invertible Weight Trajectory Decomposition |
| 532 | PE-Mamba: Bidirectional Selective Layer Aggregation for AI-Generated Image Detection |
| 541 | TAP-Out: Tracking Any Point in 360 via Out-painting |
| 542 | MixSIS: Mixed-supervised instance segmentation |
| 543 | PhysBeam: Physics-Informed Beam Splatting for Snowy LiDAR Simulation |
| 549 | Beyond False Stability: High-Noise Drift Gating for Test-Time Adversarial Defenses in Vision-Language Models |
| 551 | CoRe-CLIP: Restoring Final-Layer Patch Coherence for Training-Free Open-Vocabulary Semantic Segmentation |
| 552 | Investigating Adversarial Robustness of Multi-modal Large Language Models |
| 555 | Part-Aware Prompt Adaptation for 3D Scene Affordance Segmentation |
| 556 | OmniSurvival: Ensemble Mixture Density Learning for Oncology Survival Prediction |
| 560 | Context-Guided Semantic Alignment for Feature Fusion Networks |
| 564 | Hessian-Guided Spatial Diversity: Boosting Adversarial Transferability via Weighted Curvature Suppression |
| 565 | Adapting Foundational Image Models to Video Action Recognition Model via Learnable Context Tokens |
| 569 | AliGen: Benchmarking and Advancing Few-Shot Industrial Anomaly Generation under Object-Level Misalignment |
| 570 | Virtual-Domain-Guided Cross-Task Domain Adaptation with Scheduled Generated Supervision |
| 572 | SlotAVS: Object-Centric Audio-Visual Segmentation via Bi-Modal Slot Attention |
| 575 | MFT: An Architecture-Agnostic Adapter for Foundation Models in Medical Segmentation |
| 579 | PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting |
| 581 | EyeTAG: Eye Trajectory-Aware Gaze Estimation |
| 582 | Hierarchical Two-Stage Multi-Modal Alignment via Unified Latent Structured Representation |
| 583 | Why Multi-Source Audio Fails to Compose: A Geometric Diagnosis and Its Remedy |
| 585 | TANGO: Logit-Normalized Distillation Preserves OOD Ability in Foundation Models and Reshapes Which Scores Work |
| 586 | CSDiffWind: A Condition-Sensitivity-Distilled Diffusion framework for Low-Altitude 3D Wind-Field Forecasting |
| 588 | Geometry-Grounded Unified 3D Perception for Autonomous Driving |
| 590 | Stain-Aware Wavelet Regularization for Instant Adversarial Purification in Histopathology |
| 596 | Conditional Diffusion for 3D CT Volume Reconstruction from 2D X-rays |
| 606 | CoVR-R: Reason-Aware Composed Video Retrieval |
| 615 | Is Single-View Mesh Reconstruction Ready for Robotics? |
| 616 | EM-Mamba: Elastic Multimodal Object Detection for Edge Devices via Mamba |
| 617 | ReFineR: Single image to dense point cloud registration at pixel level with point splatting |
| 624 | Prior-Guided Implicit Neural Representations for Single-Subject Diffusion MRI Super-Resolution |
| 627 | Deep Multimodal Object Detection via Spatial Mask Interaction and Channel Competition |
| 629 | Paired Geometric Supervision for Generalizable Deepfake Detection |
| 635 | InspectVQA: Expert-Verified Visual Reasoning for Underwater Pipe Inspection |
| 639 | Beyond Parameter Space: NTK-Guided Personalized Aggregation for Robust Federated Learning |
| 641 | Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation |
| 642 | GEKey: Learning Geodesic Eccentric 3D Keypoints via Self-Supervision |
| 652 | HNH40K: A Robust Dataset and Risk-Weighted Learning for Hand Filtering |
| 659 | Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging |
| 660 | Flow-Guided Temporal Attention for Diffusion-Based Video Frame Interpolation |
| 662 | DocAttriBench: Benchmarking Answer Grounding in Document Visual Question Answering |
| 664 | Background-Free Objectness Learning for Class-Agnostic Detection |
| 667 | MEJA: A Self-Supervised Joint-Embedding Predictive Architecture for 3D Mesh Semantic Segmentation |
| 669 | CAPE-JEPA — Cross-modal Adversarial Probabilistic Embedding JEPA |
| 680 | RGFVR: Reference-Guided Face Video Restoration with Flow Matching |
| 683 | DMPT: Distributional Multi-Prompt Tuning for Robust CLIP Adaptation under Limited Supervision |
| 690 | RL-TTT: Reinforcement Learning-Based Token Selection for 3D Test-Time Training |
| 691 | Not All Patches Are Equally Forgettable: Spatially Localized Domain Unlearning in Vision-Language Models |
| 692 | Semantic Slots for Video Object-Centric Learning |
| 700 | Semantic-Aware Structural Enhancement for Unposed Multi-View Panoramic Layout Estimation |
| 701 | Adversarially Robust Few-Shot Anomaly Detection with Vision Foundation Models |
| 703 | AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models |
| 704 | Real-time Unsupervised Object Discovery from Asynchronous Event Streams |
| 707 | When More Foundation Models Means Less: Diagnosing and Addressing Multi-View Fusion Failure |
| 708 | UniFMamba: Uni-directional Visual Mamba with Full-Feature Parallel Mixing |
| 712 | Temperature-Adaptive Transformed Teacher Matching |
| 715 | MEIA: Reliability-Aware Multimodal Learning for Joint Emotion, Intention, and Action Recognition |
| 719 | Re-calibrated Contrastive Loss for Semantic-Aligned Augmentation in Vision-Language Models |
| 720 | TESS: Transitive Knowledge Editing via Bypass First-Token Penalty |
| 722 | IESC-4DGS: Interpretable and Self-Corrective 4D Gaussian Splatting for Dynamic Street-Scene Reconstruction |
| 726 | SWARD: Stochastic Window-Attention-Based Relational Distillation for Cross-Architectural Semantic Segmentation |
| 731 | CompSplat: Compression-aware 3D Gaussian Splatting for Real-world Video |
| 739 | Statistical Prior-guided Dense Assembly for Optimization-free Object Detection Dataset Distillation |
| 740 | LagrangeGS: Non-Conservative Lagrangian System on Dynamic 3D Gaussian Splatting |
| 744 | Mind the Approximation: Fisher-Weighted SVD Compression for Vision Transformers |
| 749 | Domain Generalization-Based Disentanglement of Subject Bias from Stress Signals for Multimodal Stress Recognition |
| 755 | DARE to Generalize: Domain-Aware Robust Experts and Reliability-Aware Decisions for Cross-Domain rPPG |
| 758 | LinearMask-GS: Stable-Mask Importance Pruning for Compact 3D Gaussian Splatting |
| 763 | EgoMaize: A First-Person Maize Instance Segmentation Benchmark under Severe Field Occlusion |
| 767 | TC-Omni: Temporally Consistent Omnidirectional Stereo Matching |
| 769 | FairReL: Deepfake Detection using Fairness-Aware Representation Learning |
| 773 | Continual Concept Erasure and Restoration of Diffusion Models |
| 774 | MAP3DNet: Multimodal Aesthetic Prediction for 3D Models |
| 780 | UniFaceTalk: Universal One-Shot 3D Talking-Head Synthesis via Motion-Disentangled Gaussian Splatting |
| 782 | Probing Association Instability with Track-State Perturbations for Clip-Level Active Learning in Query-Propagation Multi-Object Tracking |
| 786 | RefineFlow : Estimating the Editing Velocity with Geometric, Trajectory, and Locality Priors for Inversion Free Flow Editing |
| 791 | Are Image Generators Zero-shot Perceivers? A Rigorous Evaluation |
| 800 | Learning Structured Angle-Aware Representations for Radar Object Detection |
| 801 | UniRes: Degradation Aware Disentangled Feature Learning for Unified Image Restoration |
| 803 | DROA-CLIPSeg: Prompt-Guided Thin Crack Segmentation in Low-Light Conditions |
| 805 | MultiFlow: Vision-Driven Multimodal fMRI Encoding via Stimulus-Informed Source Flow Matching |
| 809 | ChromaIter: Real-Time Low-Light Enhancement via Iterative Log-Domain and Diversity-Enforced Reparameterization |
| 812 | Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detection |
| 825 | PhasorNet: Learning Structure from Frequency for Real-Time Stereo Matching |
| 828 | TamperLens: Tool-Augmented Vision-Language Agents for Document Forgery Detection and Localization |
| 832 | Oracle Bone Script Recognition with Efficient Topo-Swin Transformer |
| 839 | From Hierarchical Backoff to Reliable Open-World Inference |
| 856 | Hierarchical, Interpretable, Label-Free Concept Bottleneck Model |
| 860 | A Smaller Transformer in Your Transformer |
| 863 | SFA-MTKD: Spatial-Frequency-Aware Multi-Teacher Distillation for Unified Image Forgery Detection and Localisation |
| 867 | NeuDonatello: Uncertainty-Aware Framework for Accurate Neural SDF Learning |
| 868 | Gating Vision–Language Graders: A Multi-Task Detector for Reliable Evaluation of Handwritten Student Work Captured In-the-Wild |
| 870 | A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision |
| 872 | ExQuery: Explicit Structured Query Priors for Temporal 3D Object Detection |
| 879 | Quality-Preserving Online Streaming Dynamic Gaussian Splatting |
| 880 | When Does Multi-Domain Transfer Help? Structured Breast Imaging Report Generation Across Four Modalities |
| 892 | PEEK: Picking Essential frames via Efficient Knowledge distillation |
| 895 | Dynamic Alignment and Calibration for Multimodal Learning |
| 899 | PASS-3D: Pose-Anchored Sphere-Ray Shuttle 3D Reconstruction from Monocular Broadcast Badminton |
| 907 | CoCam4D: Geometry-Aware Cooperative 4D Perception for Camera-Only Autonomous Driving |
| 912 | CheXGround: Anatomical Region Tokens for Grounded Longitudinal Chest X-ray Interpretation |
| 914 | MOTIP2: Spatial Priors for End-to-End Multi-Object Tracking |
| 918 | CAR-CVGL: Conditional height-Aware BEV Representation for Cross-View Geo-Localization |
| 920 | Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association |
| 928 | RCFormer: Reliability-aware Context Transformer for Occlusion-Robust Facial Landmark Detection |
| 930 | RAIDAL: Redundancy-Aware Information Density Active Learning for CTC-Based Continuous Sign Language Recognition |
| 931 | Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion |
| 939 | Dense Indoor 3D Scene Recovery from Radar via Geometric Prior Distillation |
| 942 | Back to The Feature: Zero-Shot 6DoF Pose Estimation via Dense Local Features |
| 944 | Unlocking Spatial Grounding in Flow-Matching Models via SNR Tuning |
| 945 | MoE-AdURA-Net: An Uncertainty-Routed Mixture-of-Experts for Selective Multi-label Chest X-ray Classification |
| 949 | Intervention-Conditioned Multimodal Temporal Self-Supervised Learning for Retinal Disease Progression |
| 950 | CRLoMTL: Conflict-Resolution through Low-Rank Matrices for Multi-Task Learning |
| 951 | HYMN: Hybrid Mamba UY-Network with Mask-Encoded Prior for Laparoscopic Smoke Removal |
| 952 | BrainBind: Bridging Synchronized EEG and fMRI through Long-Form Naturalistic Video Representations |
| 962 | Stochastic Nonlinearities Improve Uncertainty Estimation |
| 967 | From Patches to Pixels: Dual-Branch Prompt Learning for Hyperspectral Scene Generalization |
| 971 | Memory Bandwidth, Not FLOPs: Profiling and Accelerating Query-Based Segmentation Decoders on Edge GPUs |
| 973 | OAEFlow: Occlusion-Aware Recurrent Encoding and Duration-Stratified Evaluation for Optical Flow |
| 974 | QTExtra: Query-Based Multi-Step Extrapolation for Asynchronous Autonomous Driving Perception |
| 981 | End-to-End Occlusion Ordered Semantic Instance Segmentation |
| 989 | RDM: Recurrent Diffusion Model for Human Motion Generation |
| 1001 | DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models |
| 1005 | GARFIELD: Graph-Adaptive SSM for Explainable 3D Multi-Person WiFi Pose Estimation |
| 1010 | DOGS: Design-Space Sampling for Prompt-Driven Logo Generation |
| 1011 | PRISSM: PRV-Guided SSM for Non-contact Stress Estimation |
| 1012 | SinoDiff: Physics-Consistent Self-Supervised Diffusion for Unified Low-Dose to Standard-Dose PET Sinogram Recovery |
| 1015 | Gated Spatial Redundancy Projection for Pathology Transformer Attentions |
| 1018 | Rethinking Test-Time Adaptation for Streaming Person Re-Identification |
| 1028 | Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgery |
| 1041 | Restoring Without Forgetting: Continual Learning Across Image Degradations |
| 1046 | CANDLE: Test-Time Debiasing of CLIP Against Typographic Attacks by Suppressing Text-Spotting Bias |
| 1047 | Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model |
| 1051 | BUSTER: Adaptive Sampling for VLM-guided Unsupervised Video Anomaly Detection |
| 1064 | Real-Time Dental Panorama Generation from Handheld Intraoral Video |
| 1065 | Dynamic Regularization for Adaptive Conformal Prediction in Deep Image Classifiers |
| 1070 | SMART: MLLM-guided Temporal Alignment for Unifying Sign Language Recognition and Spotting |
| 1074 | Lighting-Aware Diffusion-based Data Augmentation for Robust Low-Light Re-Identification |
| 1076 | CG-GLORE: A Conjugate Gradient–Based Global-Local Regularization Network for Sparse-View CT Reconstruction |
| 1078 | Beyond Calibration : Improving Mixup for Medical Image Classification via Feature-Space Selective Confidence |
| 1086 | Advancing Semiconductor Inspection: A New Dataset and Approach for Robust Anomaly Detection with Large Vision Language Models |
| 1087 | Class-Conditional Closed-Form Low-Rank Merging for Federated Continual Learning |
| 1105 | Sequence Models as Proxy for Flow Matching Based Alignment |
| 1119 | RGCT: Region-Gated Competitive Transport for Training-Free Cross-Domain Few-Shot Recognition |
| 1120 | $\text{DA}^2$: Dataset-Aware Adaptive Augmentation for Few-Shot Class-Incremental Learning |
| 1123 | Mutual Evidence Transport for Training-Free Few-Shot Recognition |
| 1131 | DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts |
| 1133 | Mars-JEPA: Multispectral Joint Embedding Predictive Architectures for Martian Landslide Segmentation |
| 1138 | Semantic-Geometric Hypothesis Verification for Cross-Platform Outdoor 3D Visual Grounding |
| 1156 | Spherocylinder Pose from Silhouette |
| 1157 | Dual-Domain Road Patch Attacks on Vision-Based 3D Lane Detectors |
| 1161 | Rethinking Microscopy Generation: Co-Designed Diffusion for Biologically Interpretable Single-Cell Synthesis |
| 1170 | Anti-Shortcut Distillation via Temporal Negative Knowledge Transfer |
| 1171 | FlowSat: Flow-Matching Diffusion Transformers with Metadata Conditioning for Satellite Image Generation |
| 1186 | Visual Counterfactual Explanations with Compositional Generative Models |
| 1187 | 3DiffGS: 3D Gaussian Splats from Unposed 2D Images using Diffusion Models |
| 1190 | SourceReward: Source-Preserving Reward Modeling for High-Precision Image Editing |
| 1194 | Distributed Semantic Segmentation With Improved Rate-Distortion Trade-Off |
| 1197 | Continuity-Driven Representation Regularization for Industrial Defect Detection |
| 1199 | Gradient-Guided Role-Aware CAM Distillation |
| 1206 | Reinforcement Learning Based Fair Adversarial Training |
| 1207 | Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation |
| 1217 | RevalExo: A Functional Daily-Activity Benchmark for Inertial and Visual Locomotion Mode Recognition in Older Adults and Clinical Cohorts |
| 1221 | PPS: Plug-and-Play Saccadic Vision for Fine-Grained Classification |
| 1230 | KGRF-Seg: Knowledge-Guided One-Step Rectified Flow Model for Medical Image Segmentation |
| 1237 | Articulate3D: Zero-Shot Text-Driven 3D Animal Asset Posing |
| 1241 | Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting |
| 1243 | BlobBoards Robust Markers for Accurate Pose |
| 1245 | TCD: Timestep-wise Contrastive Denoising for Condition-Faithful Diffusion |
| 1252 | Physics-Informed Modeling for Wood Thermal Analysis and Prediction |
| 1256 | PERSIST: Persistent-State Discrimination for Shot Boundary Detection |
| 1258 | GeM-NR: Geometry-Aware Multi-View Editing for Nonrigid Scene Changes |
| 1279 | Paraphrase Robustness in Fine-Grained CLIP: A Joint Visual–Lexical Failure Mode on Cluster-Member Classes and a Cluster-Aware Soft-Prompt Recovery |
| 1280 | RafeVPR: UAV Visual Localization with Adaptive Region Partitioning and Fourier Residual Enhancement |
| 1281 | Curriculum-guided Change Detection Training: Toward Accurate Serac Fall Monitoring |
| 1284 | LiDSeg: Linear DehazeFormer Guidance for Efficient Semantic Segmentation in Adverse Weather Conditions |
| 1288 | ActiveAugment: Online Active Learning for Augmentation Selection in Deep Learning |
| 1291 | Beyond Report Imitation: Clinically Aware Multi-Image Ultrasound Report Generation from Visible Evidence |
| 1292 | Sparse Competition during Training For the Emergence of Specialized Modules |
| 1293 | T-Time: Test-time Merging for Domain-Incremental Learning |
| 1295 | QINA: Quantum-Inspired Nonlinear Adapters for Pretrained Vision Models |
| 1296 | LiteUAV-DETR: Scale-Aware Feature Routing for Real-Time UAV Detection |
| 1297 | GaitPPT: Parallel Part-based Transformer for Gait Recognition from Lidar Point Clouds |
| 1303 | Inverse-and-Edit: Simple and Effective Framework for Fast Image Editing |
| 1309 | Lightweight Interpretable RGB-Guided Hyperspectral Super-Resolution under Real Cross-resolution Misalignment |
| 1321 | CAFIL: Concept-Aware Feature Invariance Learning for Annotation-Free Spurious Correlation Mitigation |
| 1324 | Test-Time Prototype Adaptation for Open-Vocabulary Semantic Segmentation |
| 1325 | SubPixR: A Generic Iterative Sub-Pixel Refiner for Point Tracking and Feature Matching |
| 1331 | Robust Utility Networks Registration with Powerline Segmentation on Noisy Point Clouds |
| 1332 | Beyond Classification: Structured Supervision Aligns Visual Evidence with Medical Semantics |
| 1336 | UnRL: Uncertainty-Aware RL-Controlled Adaptive 3D Mapping |
| 1343 | Simplex Diffusion for Ordinal Facial Action Unit Intensity Estimation |
| 1347 | Image Classifiers are Efficient Self-Supervised Video Representation Learners |
| 1351 | DrivingVoxels: Compositional Sparse Voxel Rasterization for Dynamic Driving Scene Reconstruction |
| 1359 | Continuous-time 4D Reconstruction via Event-guided Latent Feature Interpolation |
| 1360 | On Structured Disentangled Representation and the Limits of Global Scores |
| 1364 | SynthFaces: Balanced Large Scale Human Dataset |
| 1385 | Mistaking Periodicity for Manipulation: JPEG-Induced Structural Bias in Document Tampering Detection |
| 1387 | Sketch2TikZ: Multi-Reward Reinforcement Learning for Converting Hand-Drawn Sketches to Executable TikZ |
| 1390 | Simplified Cross-Modal Calibration for Heterogeneous Event-RGB Stereo Systems |
| 1397 | Beyond Pseudo-Masks: Semantic Disentanglement for Weakly Supervised Car Damage Segmentation |
| 1400 | Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots |
| 1411 | RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models |
| 1413 | Motion-Equivariant Pseudo-Video Augmentation: Video-Style Generalization from a Single Static Polyp Dataset |
| 1419 | FetAngle: Toward Generalizable Automated Fetal Brain Angle Biometry via Test-Time Adaptation |
| 1437 | ETNA: EnTropy regularization for compressed Neural Avatar |
| 1445 | Gated Tabular-Conditioned Attention for Robust Multimodal Radiogenomic Glioblastoma Outcome Prediction |
| 1448 | HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation |
| 1469 | BasketEvent: Understanding Who Did What and When in Basketball Videos |
| 1475 | StyleAT: Defending Face Recognition Against Semantic Attacks |
| 1497 | LACE: Lag-1 Autocorrelation-based Channel Excitation - A Plug-and-Play Module for Corruption-Agnostic CNNs |
| 1512 | SynGlass: A Large-Scale Synthetic Dataset for Fine-Grained Eyeglasses Segmentation |
| 1513 | MoTE: Mixture of Task Experts for Multi-Task Video Understanding |
| 1514 | Cross-Domain Tracker Adaptation Without Target-Domain Labels via Vision-Language Agents |
| 1516 | Towards Reliable AI-Based Histological Staining: A Systematic Study of Scaling and Uncertainty in Unpaired Generative Models |
| 1523 | Learning Ellipsoid–Ellipse Geometry for 6D Object Pose from RGB and Object Size |
| 1531 | GATE: Reliability-Gated Gaussian Evidence Fusion for Training-Free Test-Time Adaptation of Vision-Language Models |
| 1536 | Selective Prior-Conditioned State-Space Model for Remote Sensing Change Detection |
| 1538 | SM4RT: Cascaded Feed-forward Model for 4D Reconstruction |
| 1548 | TASP: Task-Agnostic Structural Pretraining for Generalizable Medical Image Segmentation |
| 1556 | Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification |
| 1561 | Bridging Asymmetric Domains: One-to-Many Image Translation for Shoeprint Retrieval |
| 1567 | Circumventing Magnitude Collapse via Modular Two-Stage Open-Set Recognition |
| 1576 | QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation |
| 1577 | FlexSplat: Flexible Feed-Forward 3D Gaussian Splatting without Point Cloud Correspondence |
| 1578 | CoVAtt - Content-Based Verification for Attribution of AI-Generated Images |
| 1583 | CrACK: Adversarial Attacks on Cross-Model Consistency in Collaborative Vision Foundation Models |
| 1590 | SlotDiT: Object-Centric Representations for Diffusion Transformers |
| 1595 | Reachability Is Not Generalization: Understanding Verb--Noun Decomposition in Assembly Action Recognition |
| 1598 | MedCORE: Criteria-Grounded Clinical Reasoning for Interpretable Medical Image Diagnosis |
| 1606 | AMGF3D: Adaptive Multi-scale Gated Fusion for Robust Indoor RGB-D 3D Object Detection |
| 1607 | X$^2$Localizer: Cross-Grained Alignment for Progressive Cross-View Video Geo-Localization |
| 1615 | Few-Shot Logical Anomaly Detection via Symbolic Constraint Extraction with Grounded Explanations |
| 1624 | A Feasibility Study on Self-Supervised LLM-Inspired Training for Generalizable Human Motion Understanding |
| 1628 | Evaluating Video LLMs’ Understanding of Corner Cases in Autonomous Driving |
| 1632 | OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting |
| 1643 | Triple-Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation |
| 1645 | ReCalMatch: Reliability-Calibrated Semantic Guidance for Semi-Supervised Fine-Grained Recognition |
| 1647 | SPDistill: Sparse Pruning-Aware Distillation for Efficient 3D Small Object Detection |
| 1651 | Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers |
| 1652 | Multi Activity Sequence Alignment via Implicit Clustering |
| 1664 | Human Video Generation from a Single Image with 3D Pose and View Control |
| 1665 | Weakly Supervised Affordance Grounding via Semantic Affordance Anchoring |
| 1666 | Provably Contractive and High-Quality Denoisers for Convergent Restoration |
| 1675 | Privacy–Utility Tradeoffs in Remote Photoplethysmography from Facial Video |
| 1676 | TetherCache: Stabilizing Long-Form Video Generation with Gated Recall and Trusted Alignment |
| 1680 | Quotient-Space Token Context Memory for Robust Visual Object Tracking |
| 1681 | SignRR: Retrieve and Refine Real Motion for Sign Language Production |
| 1701 | Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models |
| 1707 | Adaptive Subspace Projection for Generative Personalization |
| 1714 | Gland-Boundary Structure-Aware Networks for Generalizable Gleason Grading |
| 1716 | Diversity-Aware Foundational Model-Guided Open-Set Active Learning for Histopathology |
| 1731 | Diagnosing the Structural Blind Zone in Subject-Driven Generation: A Component-Level Diagnostic Framework |
| 1736 | R2Fusion: Reliability-Aware Recurrent Fusion for Burst Image Super-Resolution |
| 1744 | NeuroScope: Causal Probing of Sparse Brain-Token Interfaces in Frozen Diffusion Models |
| 1747 | MobileGestMesh: Pretrain-Free 3D Hand Mesh Estimation via Multi-Scale Token Mixer and Bounded Two-Stage Decoding |
| 1760 | DISASeg: Bridging DINO and SAM via Frequency-Aware Fusion for Automated Few-Shot Medical Image Segmentation |
| 1772 | Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study |
| 1775 | Spectral Mutual Inhibition Guided Coupled Tensor Dictionary Learning for CT–MRI Image Fusion |
| 1783 | DUIS: Detection Under Incomplete Sensing for Driving Scenes |
| 1786 | Repurposing Motion Priors: VFI-Guided Task-Agnostic Temporal Consistency Correction |
| 1794 | DoublesEval: Diagnosing Multi-Agent Tactical Reasoning in Vision-Language Models via Professional Doubles Badminton |
| 1796 | NullGuard: Null-Space Embedding for Driftless Invisible Image Watermarking |
| 1809 | Survival-Based Collision Anticipation with an Interpretable Kinematic Memory |
| 1816 | BATS-Net: Boundary-Aware Token Selection for Text-Guided Ultrasound Image Segmentation |
| 1821 | Improving Certified Robustness via Adversarial Distillation |
| 1828 | Orthogonal Polynomial Approximation for Matrix Log Normalization in Global Covariance Pooling |
| 1830 | GRACE: Generalizable Residual Autoencoder for Cross-subject EEG Decoding |
| 1843 | Advancing Trademark Image Retrieval: A Real-World Benchmark and Foundation Model Adaptation |
| 1845 | Seeing Step-by-Step: Progressive Semantic Feedback for Joint Fusion and Segmentation |
| 1853 | Physically-Grounded Spring-Energy Constraints for Fine-Grained Vision-Language Captioning |
| 1856 | Food Image Nutrition Estimation with Open-Weight Large Multimodal Models: Fine-Tuning, Multi-View Ensembles, and Retrieval Augmentation |
| 1857 | FTR-Net: Frequency--Topology--Refinement for Stage-Aware Thin-Structure Segmentation |
| 1868 | When Neural Collapse Fails to Route: Routing Drift in Continual Learning |