Building Darb: A Dual-Engine AI Architecture for Real-Time Assistive Vision
Have you ever tried walking around your room with your eyes closed? Now imagine relying on a cloud AI model to guide you. If it takes 2 seconds to process a video frame and tell you there's a wall ahead... you've already
Have you ever tried walking around your room with your eyes closed? Now imagine relying on a cloud AI model to guide you. If it takes 2 seconds to process a video frame and tell you there's a wall ahead... you've already walked into it.
When it comes to assistive tech, latency isn't just an annoyanceβitβs a safety hazard. But if you rely only on offline, on-device models, you lose out on the insane contextual awareness that modern LLMs give you.
I ran straight into this problem while building Darb, an assistive navigation platform for the visually impaired (which ended up taking home an Outstanding Completion Award at Kuwait Codes 2026!).
My solution? Stop making one AI do everything. I split Darb's "brain" into two engines.
"Reflexes" vs. "Comprehension"
Think about human biology. If someone throws a ball at your face, you duck instantly (a reflex). You don't consciously analyze the ball's trajectory first. But when you walk into a new coffee shop, you take a second to scan the room and figure out where the empty tables are (comprehension).
Darb mimics this using two separate inference engines:
- The "Reflex" Engine (Edge ML) For immediate physical safety, Darb uses MediaPipe and local TensorFlow Lite models running right on the device.
The Job: Catch immediate obstacles, barriers, or drops.
The Speed: 30+ FPS. Zero latency.
The Best Part: It works completely offline. No waiting on network requests when you're about to trip.
- The "Comprehension" Engine (Cloud AI) For deep environmental awareness, Darb sends keyframes to Google Vertex AI (Gemini).
The Job: Semantic scene understanding.
The Speed: ~1 to 3 seconds (depending on the network).
The Output: Rich audio descriptions like, "You are in a hallway. There is an open door on your right, and a person walking towards you."
The Stack That Makes It Work
Core: React and Next.js (fast, snappy, and scalable).
Vision: MediaPipe (for the zero-latency client-side spatial mapping).
Cloud: Vertex AI (for the heavy-lifting semantic analysis).
The Glue: I wrote custom state-management hooks so the "Reflex" engine can instantly interrupt and override the "Comprehension" engine if it spots a hazard while the cloud is still "thinking."
The Biggest Lesson
Building Darb taught me that AI engineering is way more than just pinging the smartest API you can find. Itβs about actual systems architecture. Itβs about knowing where the computing should happen based on the real-world physical constraints of the person using your app.
By mixing the raw speed of Edge ML with the depth of Vertex AI, I didn't have to choose between keeping users safe and keeping them informed.
Have you guys messed around with hybrid edge/cloud ML setups? I'd love to hear how you handle the latency vs. context trade-off!
Abdulrahman Haramain is a software engineer and AI systems developer based in Kuwait. He builds production web platforms and offline computer-vision applications. Connect on GitHub, LinkedIn, or at abdulrahmanharamain.com.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.