Caner Olcay
All projects

NeuroPark — 3D vehicle detection

YOLOv8, MiDaS depth, and Open3D turning a single traffic-camera image into an oriented 3D vehicle scene.

Role
Team Lead / ML Engineer
Timeline
2025
Status
Presented at KPZ25
Area
Computer vision

Overview

A computer vision project run with Neurosoft that explores monocular perception for 3D scene understanding. The pipeline combines object detection, depth estimation, point-cloud generation, and spatial alignment to produce usable vehicle scenes without relying on LiDAR as the primary sensor.

Problem

3D scene understanding helps with autonomous parking and traffic analysis, but richer sensors are expensive and not always available.

Approach

A pipeline that takes monocular RGB input, estimates depth, reconstructs point clouds, and localizes vehicles in aligned 3D world coordinates.

How it works
  1. 01

    Detect

    YOLOv8

  2. 02

    Estimate depth

    MiDaS

  3. 03

    Reconstruct

    Open3D point cloud

  4. 04

    Fit boxes

    PCA / SVD geometry

  5. 05

    Validate

    Against LiDAR

YOLOv8 handles detection, MiDaS estimates depth, Open3D builds colored point clouds, and PCA/SVD-based geometry fits oriented 3D boxes aligned with world coordinates.

Outcomes

  • Led a 3-person team, together with Neurosoft: one 2D traffic-camera image in, a 3D vehicle scene out.
  • Combined YOLOv8 detection, MiDaS depth estimation, and Open3D, checked against LiDAR measurements.
  • Reached roughly 5–6 s per high-resolution image and presented the system at the KPZ25 Engineering Conference.

Challenges

  • Depth uncertainty from a single camera.
  • Spatial alignment and coordinate calibration across the whole pipeline.
  • Balancing reconstruction quality against processing time.

Stack

YOLOv8 · MiDaS · OpenCV · PyTorch · Open3D · SciPy