CVPR 2026 目标检测(object detection)方向上接收论文总结2

接前文:CVPR 2026 目标检测(object detection)方向上接收论文总结1

其他

以下论文标题中出现 detect/detection 等字样,但不完全属于目标检测主线(如异常检测、伪造/生成内容检测、OOD 检测、变化检测、动作检测等)。按子方向分组列出。

异常检测

  1. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection
  1. ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning
  1. Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detection
  1. Anomaly-Related Residual Fields for Cross-domain Anomaly Detection
  1. AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
  1. Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
  1. Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
  1. CHAL: Causal-guided Hierarchical Anomaly-aware Learning for Moving Infrared Small Target Detection
  1. Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection
  1. Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection
  1. DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual Prompting
  1. Dual-Prototype-Guided Multi-task Learning for Unsupervised Anomaly Detection and Classification
  1. FastRef: Fast Prototype Refinement for Few-shot Industrial Anomaly Detection
  1. FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
  1. Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity Learning
  1. From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection
  1. Geometry-Aligned and Anomaly-Aware Reconstruction for 3D Anomaly Detection
  1. GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection
  1. GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning
  1. Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
  1. Hunting Normality from Query Sample via Residual Learning for Generalist Anomaly Detection
  1. InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
  1. Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection
  1. LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly Detection
  1. Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly Detection
  1. MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
  1. MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection
  1. Multi-Prototype Compactness and Boundary-Aware Synthesis for Unsupervised Anomaly Detection
  1. No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
  1. Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly Detection
  1. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection
  1. RAID: Retrieval-Augmented Anomaly Detection
  1. RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation
  1. Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision
  1. SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling
  1. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection
  1. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection
  1. Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective
  1. UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
  1. VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer
  1. Wavelet-Driven 3D Anomaly Detection under Pose-Agnostic and Sparse-View
  1. Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning

伪造/生成内容检测

  1. A Debiased Reconstruction-based Framework for Training-Free Detection of AI-Generated Images
  1. A Difference-in-Difference Approach to Detecting AI-Generated Images
  1. A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real World
  1. Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
  1. All in One: Unifying Deepfake Detection, Tampering Localization, and Source Tracing with a Robust Landmark-Identity Watermark
  1. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
  1. Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detection
  1. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
  1. Cross-modal Representation Learning for Diffusion-generated Image Detection
  1. Decoupling Bias, Aligning Distributions: Synergistic Fairness Optimization for Deepfake Detection
  1. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detection
  1. Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification
  1. Detecting Compressed AI-Generated Images via Phase Spectrum Robustness
  1. DFD-HR: Generalizable Deepfake Detection via Hierarchical Routing Learning
  1. DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
  1. Diversity over Uniformity: Rethinking Representation in Generated Image Detection
  1. Enabling Supervised Learning of Generative Signatures for Generalized AI-Generated Images Detection
  1. FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models
  1. Investigating Self-Supervised Representations for Audio-Visual Deepfake Detection
  1. Layer Consistency Matters: Elegant Latent Transition Discrepancy for Generalizable Synthetic Image Detection
  1. Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
  1. Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
  1. Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
  1. PPM-CLIP: Probabilistic Prompt Modeling for Generalizable AI-Generated Image Detection
  1. ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation
  1. SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual Learning
  1. Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
  1. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
  1. Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
  1. TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
  1. Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
  1. UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
  1. Unleashing Vision-Language Semantics for Deepfake Video Detection
  1. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation
  1. X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
  1. Your One-Stop Solution for AI-Generated Video Detection
  1. Zero-shot Detection of AI-Generated Image via RAW-RGB Alignment

OOD检测

  1. Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models
  1. ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
  1. Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal Transport
  1. Enhancing Out-of-Distribution Detection with Extended Logit Normalization
  1. Learning Latent Concepts for Detecting Out-of-Distribution Objects
  1. Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs
  1. Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis
  1. Neural Distribution Prior for LiDAR Out-of-Distribution Detection
  1. RankOOD - Class Ranking-based Out-of-Distribution Detection
  1. Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection
  1. The Invisible Gorilla Effect in Out-of-distribution Detection
  1. TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models
  1. UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language Modeling

变化检测

  1. Changes in Real Time: Online Scene Change Detection with Multi-View Fusion
  1. OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery
  1. RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection
  1. SRGCD: Stability-Driven Region Growth Framework for 3D Change Detection
  1. UniChange: Unifying Change Detection with Multimodal Large Language Model

动作检测

  1. Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
  1. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
  1. MoVie: Broaden Your Views with Human Motion for Action Detection
  1. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
  1. Streamlined Open-Vocabulary Human-Object Interaction Detection
  1. TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

关键点/地标检测

  1. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images
  1. EV-CGNet: Co-visible Focused 3D-guided 2D Event Keypoint Detection Network
  1. From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

幻觉检测

  1. Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
  1. Lyapunov Probes for Hallucination Detection in Large Foundation Models
  1. PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models
  1. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
  1. ZINA: Multimodal Fine-grained Hallucination Detection and Editing

讽刺/语义检测

  1. MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection

跟踪相关

  1. From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking

其他检测相关

  1. Adaptive Confidence Regularization for Multimodal Failure Detection
  1. ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
  1. AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Models
  1. AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Models
  1. BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
  1. Breaking Spurious Correlations: Uncertainty-Driven Causal Transformers for AU Detection
  1. Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial Transcriptomics
  1. BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship Detection
  1. Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
  1. COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learning
  1. CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
  1. Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets
  1. DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
  1. Detect Any AI-Counterfeited Text Image
  1. Detect Anything via Next Point Prediction
  1. DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imaging
  1. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection
  1. FedSDR: Federated Graph Learning with Structural Noise Detection and Reconstruction
  1. Geometry-driven OOD Detectors Are Class-Incremental Learners
  1. Ghost-FWL: A Large-Scale Full-Waveform LiDAR Dataset for Ghost Detection and Removal
  1. GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
  1. Homaloidal parametrization for detecting critical two-view configurations
  1. KLIP: Localized Distribution Shift Detection via KL-Divergence with Diffusion Priors in Inverse Problems
  1. Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
  1. Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection
  1. LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
  1. Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
  1. MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision
  1. MEMO: Human-like Crisp Edge Detection Using Masked Edge Prediction
  1. MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
  1. Neural Field-Based 3D Surface Reconstruction of Microstructures from Multi-Detector Signals in Scanning Electron Microscopy
  1. Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting
  1. OpenFS: Multi-Hand-Capable Fingerspelling Recognition with Implicit Signing-Hand Detection and Frame-Wise Letter-Conditioned Synthesis
  1. OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
  1. Physical Adversarial Clothing Evades Visible-Thermal Detectors via Non-Overlapping RGB-T Pattern
  1. Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
  1. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction
  1. ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection
  1. RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D Detection
  1. SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
  1. Scene Reconstruction as Mapping Priors for 3D Detection
  1. Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective
  1. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection
  1. Similarity-Consistent Likelihood Diffusion enables Hidden Person Detection from Wall Reflections
  1. SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images
  1. Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
  1. Target-Aware Invertible Encoder with Reconstruction Guidance for Infrared Small Target Detection
  1. Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach
  1. Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods
  1. TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
  1. TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramas
  1. UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection
  1. Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectors
  1. Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

总结

从本届接收论文来看,CVPR 2026 目标检测方向呈现以下趋势:

  1. 3D 目标检测体量最大:单目、多视角、BEV、LiDAR/Radar 融合与室内外统一检测持续活跃;雷达-相机融合、Gaussian Splatting 先验、token 压缩与不确定性估计是常见技术点。

  2. 开放词汇/开放世界检测成为主线之一:Open-Vocabulary Detection、Open-World Detection、未知类别发现、检索式检测(如 WeDetect)与热成像开放词汇检测等方向快速增长。

  3. 数据高效学习受重视:少样本、跨域少样本、增量检测、主动学习、在线数据筛选与弱监督设定显著增多,反映标注成本与持续部署需求。

  4. 实时高效架构回潮:YOLO 体系、Mamba/SSM 混合结构、轻量化模型与训练策略优化重新成为焦点。

  5. 场景专用化加深:遥感/旋转框、UAV、小目标、伪装/显著性、水下、X-ray 安检、Person Search 等方法继续细分。

  6. 检测概念外延明显:异常检测、深度伪造/生成内容检测、OOD 检测、变化检测等“泛检测”任务数量可观,但与经典目标检测主线有所区分。

总体而言,CVPR 2026 目标检测研究在通用检测框架演进之外,更强调开放词汇泛化、三维感知、数据高效学习与真实场景鲁棒落地。

参考资料

  1. CVPR 2026 Official Website

  2. CVPR 2026 Accepted Papers (Open Access)

  3. arXiv.org

(注:文档部分内容由 AI 生成;Code/Blog/单位信息以公开网页检索为准,如有遗漏欢迎补充指正。)

©著作权归作者所有,转载或内容合作请联系作者
【社区内容提示】社区部分内容疑似由AI辅助生成,浏览时请结合常识与多方信息审慎甄别。
平台声明:文章内容(如有图片或视频亦包括在内)由作者上传并发布,文章内容仅代表作者本人观点,简书系信息发布平台,仅提供信息存储服务。

友情链接更多精彩内容