CV Brief · Wednesday, 17 June 2026
CV Brief
Research & Papers
RAMS: Dynamic Model Switching for Embedded YOLOv8 Edge Detection
RAMS is a lightweight runtime controller that dynamically switches between YOLOv8 NANO/SMALL/MEDIUM tiers based on device resource pressure, without model-reload latency. It monitors CPU/memory constraints and calibrates switching thresholds from idle behavior, solving the real production problem of balancing inference speed and accuracy on resource-constrained hardware.
Read more →VigilFormer: Deformable Attention for Real-Time Video Anomaly Detection
VigilFormer combines deformable spatio-temporal attention with causal temporal modeling to detect anomalies in untrimmed surveillance video while maintaining real-time throughput. Directly addresses the production tension between detection accuracy and inference speed in surveillance systems.
Read more →CNN vs Vision Transformers for Maritime Object Detection Under Adverse Weather
Comparative evaluation of CNN and ViT architectures for ship detection across 6,468 maritime images covering cloudy, foggy, and rainy conditions. Provides empirical guidance on architecture selection for real-world maritime surveillance where weather robustness directly impacts deployment decisions.
Read more →Tools & Releases
Track Class Lock fixes video label flickering in real-time
Roboflow introduces Track Class Lock, a workflow block that stabilizes object class predictions across video frames by freezing labels once detector confidence settles. Solves a common production pain point where detection models rapidly flip between classes on the same tracked object.
Read more →IV bag detection combines RF-DETR with multimodal LLM analysis
Roboflow demonstrates industrial CV pipeline using RF-DETR object detection plus Gemini 2.5 Pro for fill-level quantification and leak classification in medical devices. Practical template for combining traditional detection with LLM reasoning in production systems.
Read more →Automated tire OCR extracts sidewall codes via vision agents
Roboflow Vision Agent combines RF-DETR detection with multimodal LLMs to extract DOT codes and tire specifications from sidewall images at scale. Shows how to build modular CV pipelines that handle both detection and structured data extraction.
Read more →Tutorials & Guides
Netflix VOID model: three failure surfaces in production CV
Deep postmortem on resolution degradation, identity tracking, and quadmask conditioning failures in Netflix's real-world vision deployment. Exposes the gap between lab performance and production reality—critical for engineers shipping video CV systems.
Read more →Seven ML infrastructure tricks: slow experiments to production scale
Practical infrastructure patterns that bridge the gap between slow prototypes and scalable production systems. Essential for CV teams managing experiment-to-deployment pipelines.
Read more →Industry & Deployments
V-INTELLIGENCE: full-stack vehicle compliance with plate tracking
End-to-end traffic management system combining video processing, plate detection, and automated enforcement. Real-world case study in building deployed CV pipelines at scale.
Read more →pHash deduplication for video crops: use Hamming distance ≤10 as your threshold. Too tight misses duplicates, too loose removes valid unique crops.
Quick Links
- Disagreement-Based Cross-Model Routing for Implicit Video Question Answering
- Interpolation between Convolution and Attention via K-Nearest Neighbors
- FairGen: Preference-Aligned Diffusion for Demographically Equitable Medical Imag
- FUSE: Quantifying Uncertainty in Vision-Language Models by Bayesian Fusing Epist