Real-Time Object Detection Latency in Robot Perception Pipelines
Latency budgets depend on the task, not industry benchmarks.

Real-time object detection latency is a budget that changes shape depending on the task. It's a budget that changes shape depending on the task, and that budget gets spent across a chain of decisions: which model runs, which chip it runs on, how many sensors feed it, and whether reasoning sits in the loop or off to the side. Understanding where the milliseconds actually go, rather than treating "real-time" as a fixed target, is what separates a perception stack that works from one that only benchmarks well. Real-Time Object Detection Latency in Robot Perception Pipelines.
Why "real-time" means different things to different robots
There's no industry-wide clock that defines real-time perception, and treating one benchmark as universal is a common mistake. The frame rate a system needs is set entirely by what the robot is doing.
Vision-language models deployed on edge devices sit at the opposite end of the range. Ten frames per second is treated as the floor for interaction to feel natural to a person standing in front of the robot, and reaching even that floor on constrained hardware turns into a real optimization problem, not a checkbox Kong et al., Scientific Reports Park, Kim & Ko, Frontiers in Robotics and AI.
The most unforgiving case comes from autonomous racing, where vehicles operate near 80 meters per second Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team. At that speed, 100 milliseconds of latency, barely enough time to blink, translates into almost 8 meters of spatial displacement Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team. In autonomous racing near 80 m/s, 100 milliseconds of latency corresponds to nearly 8 meters of spatial displacement, making even small delays mission-critical Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team. That gap is the whole argument for treating latency as task-defined: the number that counts as acceptable in one deployment would be disqualifying in another. Before any model or chip gets chosen, the latency budget the task actually demands has to be pinned down first, because everything downstream inherits that number. Surveillance and consumer video pipelines typically operate at 25–30 FPS (Kong et al., Scientific Reports 16, 5875, 2026). High-speed tasks such as UAV obstacle avoidance require 50–60 FPS (ibid.) Kong et al., Scientific Reports.
The full pipeline: all the places latency accumulates before an output appears
End-to-end latency covers the full wall-clock span from the moment a robot receives an image, and a text prompt if the pipeline uses one, to the moment it outputs a finished mask or bounding box. That includes preprocessing, inference, and postprocessing, not inference time alone Kong et al., Scientific Reports. Preprocessing covers voxelization, bird's-eye-view rasterization, resolution downscaling, and tokenization, work that happens before the network ever sees a usable tensor. Postprocessing carries its own tax too: non-maximum suppression cleans up overlapping detections, but it adds latency and depends on IoU thresholds tuned by hand, which makes it a fragile step to carry into production Kong et al., Scientific Reports.
Multi-sensor timing hides an entire category of cost that inference benchmarks never show. Cameras typically run at 30 to 60 Hz, LiDAR at 10 to 20 Hz, and radar somewhere around 10 to 30 Hz, and because those streams arrive on different clocks, a fusion pipeline has to align timestamps, buffer the slower sensors, and interpolate between measurements, with each of those steps adding delay before fusion even starts Kong et al., Scientific Reports Park, Kim & Ko, Frontiers in Robotics and AI.
A per-stage breakdown from a robotic manipulator running on an RTX 3080 laptop GPU makes the point concrete. Keypoint detection cost 63.7 milliseconds, by far the largest share of the budget arxiv.org. Inpainting reconstruction added 21.2 milliseconds, while the outlier gate and UKF update together stayed under 2.2 milliseconds arxiv.org. The full pipeline with inpainting ran at 87.1 milliseconds, or 11.5 Hz; stripped of inpainting, it dropped to 65.9 milliseconds, or 15.2 Hz, and both versions fit inside the 100-millisecond budget the 10 Hz visual servoing loop demanded Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team arxiv.org. Keypoint detection, not the control math, ate the budget, a fact that appears only when a team profiles individual stages rather than "inference time" in the aggregate.
Resolution deserves to be treated as a design parameter in its own right rather than a setting left at default. Work on resolution-aware parameterization shows that modeling latency as a function of input resolution gives a measurably better picture of how a pipeline actually behaves under load, which makes resolution a lever, not an afterthought.
Model architecture choices for the speed–accuracy dial in closed-set detection
For closed-set detection, where the list of object categories is fixed ahead of time, YOLO remains the model family most robotics teams reach for first, and the line has kept moving toward faster, simpler deployment with each generation.
The published numbers back the design up. For robotics specifically, that reduction in inference delay improves tasks like robotic arm grasping under dynamic conditions and mobile robot obstacle recognition in cluttered environments. The model has been benchmarked against YOLOv8, YOLOv11, YOLOv12, YOLOv13, RF-DETR, and RT-DETR on Jetson Nano and Orin hardware, and it exports to ONNX, TensorRT, CoreML, and LiteRT with INT8 quantization; FP16 isn't a separate export format, it just runs automatically at inference time through the GPU delegate.
RF-DETR, released under Apache 2.0, takes the opposite bet: accuracy and generalization over raw speed Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team. Roboflow's analysis puts RF-DETR ahead of YOLO26 on accuracy benchmarks and on domain transfer, measured against the RF100-VL real-world transfer benchmark, and the model runs on edge hardware through Inference Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team. It's the better fit when a project needs to generalize across novel domains more than it needs the lowest possible latency Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team. Neither model is wrong for the wrong reason; they're solving for different constraints.
Dropping NMS isn't free, either Kong et al., Scientific Reports. An NMS-free architecture changes the training regime a team is working within, so the decision to adopt one comes down to weighing deployment simplicity against however much calibration effort the existing workflow already tolerates. Even the largest YOLO26 scale leaves substantial headroom within a 100-millisecond budget, but that headroom gets consumed by the preprocessing and fusion stages described earlier Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team Sapkota et al., YOLO26. The model is rarely the bottleneck; the pipeline around it usually is Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team Sapkota et al., YOLO26. Key architectural changes enabling this shift include removal of Distribution Focal Loss (DFL), end-to-end NMS-free inference, ProgLoss and Small-Target-Aware Label Assignment (STAL), and the MuSGD optimizer (Sapkota et al., arXiv:2509.25164, 2026) Kong et al., Scientific Reports. Across five detection scales, benchmarks span 40.9–57.5 mAP on COCO at 1.7–11.8 ms T4 TensorRT latency (Ultralytics YOLO26 Docs) Sapkota et al., YOLO26. On CPU inference, the model is up to 43% faster than YOLO11n on an Intel Xeon @ 2.00 GHz (ibid.) Sapkota et al., YOLO26.
Where open-vocabulary models fit and their latency cost
Vision-language models are entering robotics pipelines because closed-set detection has a hard ceiling: it only recognizes what it was trained to recognize. Zero-shot, open-vocabulary recognition lets a robot handle objects it has never seen and follow natural-language commands without retraining, which matters most in domestic settings and collaborative industrial floors where the object list can't be fully specified in advance Kong et al., Scientific Reports. The cost of that flexibility is computational, and it runs straight into the constraints of edge hardware. Hitting even 10 FPS on a resource-limited device is a genuine optimization challenge Park, Kim & Ko, Frontiers in Robotics and AI.
The result exposes a real trade-off rather than a clean winner. TensorRT-accelerated NanoOWL runs fastest but only recognizes noun phrases, while PyTorch-based YOLO-World runs slower yet understands full, complex sentences.
For pipelines where on-device inference just isn't feasible, edge-based mobile edge computing offers another path.
The deeper lesson sits in the NanoOWL versus YOLO-World comparison: the language understanding that makes YOLO-World useful for human-robot interaction is the same property blocking it from full TensorRT optimization. A team doesn't get both properties at once; they pick which constraint to relax. And the fact that NanoOWL's speed comes almost entirely from how it's compiled says something larger: model choice and hardware optimization are the same decision, made twice. They're the same decision, made twice. The YOLO-World-X model registered 45.59 ms per frame on that platform (Park et al., Frontiers in Robotics and AI). The best overall pipeline, NanoOWL + EfficientViT-SAM, derives its efficiency advantage from TensorRT optimization of the full detection engine (Park et al., Frontiers in Robotics and AI, 2025) Kong et al., Scientific Reports. In a large-scale VLM deployment via edge MEC, LLaMA-3.2-11B-Vision-Instruct over WebRTC/ORAN reduced end-to-end latency by 5% versus cloud while preserving near-cloud accuracy, and the compact Qwen2-VL-2B-Instruct achieved sub-second responsiveness, a useful framing for pipelines where on-device inference is not feasible (arXiv:2601.14921v1, January 2026) arxiv.org.
Hardware pairing and optimization techniques that move latency independently of model choice
TensorRT is the single largest lever available on NVIDIA hardware, independent of which model sits on top of it. It compiles a kernel-fused engine for the exact target GPU, and the typical payoff is a substantial multiple in speedup with the lowest inference latency a team is likely to reach on that hardware. The cost is that the engine becomes device-specific, and quantization has to be calibrated for it.
Quantization itself comes in tiers, and each one trades precision for speed differently. FP16 is the standard first step, delivering a large speedup for almost no accuracy loss. INT8 roughly doubles FP16's speedup on top of that, but it demands calibration against representative images to pick quantization ranges that don't erode accuracy. Further out, Jetson Thor's Blackwell-class hardware exposes FP4, and TensorRT now supports INT8, INT4, FP8, and FP4 together, options that matter most in next-generation edge deployments where memory bandwidth, not raw compute, is what's actually limiting the system.
None of this comes for free at every batch size, though. A 2026 benchmark on Jetson Orin NX running YOLOv8 found TensorRT beating PyTorch by 17.7% at batch size 2, but at batch size 8, PyTorch held up more consistently as TensorRT ran into its own resource limits MDPI Computers. Batch size, in other words, isn't a dial a team can crank up assuming linear gains.
Compute doesn't have to live in one place, either. NXP i.MX processors can act as distributed edge-perception nodes inside a humanoid robot, handling time-critical sensor preprocessing and control tasks alongside a centralized NVIDIA compute stack, with NVIDIA's Holoscan Sensor Bridge enabling that split. A reproducible benchmark across Rockchip SoCs found that inference latency tracked more closely with detection accuracy than with FLOPs or parameter count, multi-core NPU scheduling delivered only marginal gains because of synchronization and shared-memory bottlenecks, and memory bandwidth turned out to be the dominant factor for robustness once the chip was multitasking. TOPS ratings are weak predictors of real-world latency, and hardware-aware model selection, paired with memory-efficient optimization, affects real-world performance more than the headline compute number stamped on the box. For latency-sensitive applications such as robotics at 100 Hz and interactive AI, TensorRT FP16 with batch size 1–4 and sequence length ≤128 is the recommended configuration for transformer inference (arXiv:2603.28708) Park, Kim & Ko, Frontiers in Robotics and AI UNIMORE Racing team.
Asynchronous reasoning and perception coexisting without violating latency budgets
Reasoning modules, LLM-based planners and ReAct agents among them, operate on timescales incompatible with perception frame rates. Wiring perception to block on a reasoning module's output is a direct way to blow through a latency budget.
A deployed example makes the shape of the fix clear. Fire and smoke detection inside that pipeline hits 94.2% mAP at 127 milliseconds of latency SafeGuard ASF: SR Agentic Humanoid Robot System for Autonomous Industrial Safety. The ReAct reasoning module runs on CPU at 850 milliseconds per decision cycle, but reasoning executes asynchronously, so perception continues real-time monitoring during deliberation, and the two loops do not block each other Kong et al., Scientific Reports arxiv.org.
That 850-millisecond figure should not be glossed over Kong et al., Scientific Reports arxiv.org. In a fast, safety-critical loop, it would be disqualifying. Here, it's described as acceptable precisely because reasoning has been decoupled from perception, and the actual principle worth taking away is this: a given latency number's acceptability depends entirely on which loop it belongs to, not on the number in isolation Kong et al., Scientific Reports arxiv.org. It also explains, in hindsight, why a large VLM can have any place in a robotics system at all: it was moved off the critical path entirely. A deployed example is SafeGuard ASF running on the Unitree G1 with a Jetson Orin. Its perception pipeline runs at 25–30 FPS aggregate, with the thermal channel as the fastest component due to lower input resolution (Kong et al., Scientific Reports).
Sensor fusion timing: the latency costs that appear only when multiple modalities combine
Fusion introduces latency sources that don't exist when a robot runs a single sensor. Cameras, LiDAR, and radar operate on different native clocks, roughly 30 to 60 Hz, 10 to 20 Hz, and 10 to 30 Hz respectively, so before a fusion algorithm can combine them, the streams need their timestamps aligned Park, Kim & Ko, Frontiers in Robotics and AI. Alignment alone isn't enough: the pipeline also has to buffer faster sensors while waiting on slower ones and interpolate across measurements that never arrive in sync, and every one of those steps adds delay to the total, regardless of how fast each sensor's own inference runs Kong et al., Scientific Reports Park, Kim & Ko, Frontiers in Robotics and AI.
At high speed, those costs stop being a rounding error and start driving outcomes directly. Near 80 meters per second, asynchronous multi-sensor timing, detection delay, and high relative velocity between objects compound each other, and the result is state estimation error measured in meters rather than centimeters, error that feeds straight into a collision-avoidance system's ability to get a correct read on the world in time to act on it. Fusion timing, in that setting, is the constraint the rest of the system has to be built around. It's the constraint the rest of the system has to be built around.
Sources
- YOLO26: Key Architectural Enhancements and Performance Benchmarking for Real-Time Object Detection
- Object detection on low-compute edge SoCs: a reproducible benchmark and deployment guidelines | Scientific Reports
- SafeGuard ASF: SR Agentic Humanoid Robot System for Autonomous Industrial Safety
- Vision-Language Models on the Edge for Real-Time Robotic Perception
- Computer Vision for Collaborative Robots in Industry 5.0: A Survey of Techniques, Gaps, and Future Directions
- frontiersin.org
- arxiv.org


