NeuroPark — 3D vehicle detection
YOLOv8, MiDaS depth, and Open3D turning a single traffic-camera image into an oriented 3D vehicle scene.
- Role
- Team Lead / ML Engineer
- Timeline
- 2025
- Status
- Presented at KPZ25
- Area
- Computer vision
Overview
A computer vision project run with Neurosoft that explores monocular perception for 3D scene understanding. The pipeline combines object detection, depth estimation, point-cloud generation, and spatial alignment to produce usable vehicle scenes without relying on LiDAR as the primary sensor.
Problem
3D scene understanding helps with autonomous parking and traffic analysis, but richer sensors are expensive and not always available.
Approach
A pipeline that takes monocular RGB input, estimates depth, reconstructs point clouds, and localizes vehicles in aligned 3D world coordinates.
01
Detect
YOLOv8
02
Estimate depth
MiDaS
03
Reconstruct
Open3D point cloud
04
Fit boxes
PCA / SVD geometry
05
Validate
Against LiDAR
YOLOv8 handles detection, MiDaS estimates depth, Open3D builds colored point clouds, and PCA/SVD-based geometry fits oriented 3D boxes aligned with world coordinates.
Outcomes
- Led a 3-person team, together with Neurosoft: one 2D traffic-camera image in, a 3D vehicle scene out.
- Combined YOLOv8 detection, MiDaS depth estimation, and Open3D, checked against LiDAR measurements.
- Reached roughly 5–6 s per high-resolution image and presented the system at the KPZ25 Engineering Conference.
Challenges
- Depth uncertainty from a single camera.
- Spatial alignment and coordinate calibration across the whole pipeline.
- Balancing reconstruction quality against processing time.
Stack
YOLOv8 · MiDaS · OpenCV · PyTorch · Open3D · SciPy