How our Top 11 HopHacks project turned noisy MediaPipe pose data into reliable rehabilitation metrics for adaptive AI planning
What if a rehabilitation exercise could understand how a patient moved, instead of only asking whether they finished? That was one of the questions behind Mendly, an AI-powered stroke rehabilitation platform my team bui
What if a rehabilitation exercise could understand how a patient moved, instead of only asking whether they finished?
That was one of the questions behind Mendly, an AI-powered stroke rehabilitation platform my team built during HopHacks.
By the end of the hackathon, Mendly became a Top 11 finalist out of 87 projects and won Best Use of DigitalOcean.
I worked mainly on the computer vision side of the project: pose tracking, joint-angle and range-of-motion measurement, valid-repetition detection, tracking-loss handling, camera reliability, and the layer that turns movement into structured performance data for our adaptive AI pipeline.
The part I found most interesting was not simply getting MediaPipe to recognize a person.
It was figuring out how to turn imperfect camera observations into data that the rest of the system could actually trust.
The pipeline looked roughly like this:
Camera
โ
Pose Detection
โ
Joint Angles / Range of Motion
โ
Reliability Checks
โ
Repetition Detection
โ
Structured Performance Data
โ
Future Adaptive Planning
That sounds straightforward on paper.
It wasn't.
Getting useful movement data from MediaPipe
We used MediaPipe Pose to estimate landmarks on the patient's body.
Instead of trying to reason directly about image pixels, we could work with points representing joints such as the shoulder, elbow, wrist, hip, knee, and ankle.
For example, if we know the positions of the shoulder, elbow, and wrist, we can calculate the angle at the elbow.
Given two vectors around a joint:
vโ = Shoulder - Elbow
vโ = Wrist - Elbow
we can calculate the angle between them with the dot product:
ฮธ = arccos((vโ ยท vโ) / (|vโ||vโ|))
Once we have joint angles over time, we can start measuring things like range of motion.
If an elbow moves from roughly:
155ยฐ โ 72ยฐ
then the observed range is approximately:
83ยฐ
This is already much more useful than simply knowing that "the arm moved."
But getting an angle is the easy part.
The harder question is whether that angle should be trusted.
A pose model is not a reliable sensor
This was probably the biggest thing I learned from the project.
At first, it is easy to think of pose estimation like this:
MediaPipe detected a wrist
โ
therefore the wrist position is correct
Real camera input is much messier.
A patient can move partly outside the frame. A joint can become occluded. Lighting can change. Landmarks can jump. The patient can turn sideways. The camera can be positioned badly.
A single bad landmark can propagate through everything after it:
bad landmark
โ
bad angle
โ
bad repetition
โ
bad performance data
โ
bad future planning context
That made reliability part of the actual product, not just an implementation detail.
Some of the most interesting bugs appeared when we asked a simple question:
What should happen when the camera does not know?
One of my favorite bugs: failed inference
Suppose the tracker sees:
Frame 1: 40ยฐ
Frame 2: 55ยฐ
Frame 3: tracking fails
What should Frame 3 become?
There are two tempting answers.
One is:
0ยฐ
The other is:
55ยฐ
Neither is correct.
If tracking failed, we do not know where the patient's arm actually was.
Treating the missing value as zero invents a movement that never happened.
Keeping the previous angle is not much better. The patient may have continued moving while the camera lost them.
So we changed the pipeline to represent failed inference as an invalid observation.
Pose succeeds
โ process the measurement
Pose fails
โ mark observation invalid
โ do not update movement state
โ do not manufacture an angle
โ do not count a repetition
The rule became:
Unknown is not zero. Unknown stays unknown.
That sounds obvious after the fact, but it fixed an important class of false movement data.
Camera geometry had another hidden problem
Another issue was caused by normalized coordinates.
MediaPipe gives landmark positions as normalized X and Y values.
But X is normalized relative to the frame width, while Y is normalized relative to the frame height.
On a 640ร480 camera:
ฮx = 0.1 โ 64 pixels
ฮy = 0.1 โ 48 pixels
Those values are both 0.1 in normalized space, but they do not represent the same physical distance in the image.
If we calculate geometry as if the normalized X and Y scales are identical, joint angles can become distorted.
We fixed this by accounting for the actual camera aspect ratio before calculating angles and ray lengths.
At the same time, we kept the original normalized landmarks for things like checking whether a joint was still inside the frame.
This was a good example of a bug that was mathematically small but affected everything built on top of it.
Counting repetitions is a state problem
Another thing that looked simple at first was repetition counting.
A naive approach might say:
Count a rep whenever the joint angle crosses a threshold.
That works until real movement starts looking like this:
155ยฐ
148ยฐ
139ยฐ
121ยฐ
97ยฐ
76ยฐ
81ยฐ
95ยฐ
119ยฐ
142ยฐ
151ยฐ
Human motion does not happen in perfectly clean steps.
The angle may hover around a threshold, move backward briefly, or contain noise.
So instead of treating each frame independently, we used a state-based approach.
A complete repetition is closer to:
Start
โ
Movement begins
โ
Required range reached
โ
Hold requirement satisfied, if needed
โ
Movement reverses
โ
Return condition reached
โ
Valid repetition
Crossing one threshold is not automatically a repetition.
A rep has a beginning, progression, target, and completion.
Failed attempts should not magically complete an exercise
During testing, we found another behavior that looked reasonable in code but felt completely wrong from the patient's perspective.
Imagine the exercise target is:
5 valid reps
and the patient does:
valid
failed
failed
valid
failed
failed
failed
They have completed only:
2 / 5 valid reps
But an earlier completion rule could still stop the exercise after enough total attempts.
That meant a patient could fail several repetitions and somehow "finish" the exercise.
We changed that.
Automatic completion now depends on the number of valid repetitions, not the total number of attempts.
validRepCount >= targetRepCount
Failed attempts can still be recorded.
They just do not count toward the target.
So if the prescription says five valid repetitions, the patient actually needs five valid repetitions.
Recorded data and trusted data are not the same thing
This distinction also became important for range-of-motion statistics.
Suppose a patient performs a repetition while the camera is barely tracking the required landmarks.
We may still want to know that an attempt happened.
But that does not mean its ROM measurement should be treated as reliable.
Earlier logic could fall back to unreliable repetitions when there were no reliable ones available.
We changed that behavior.
A low-confidence attempt can remain part of the exercise history, but unreliable motion should not contaminate trusted ROM statistics.
I started thinking about this as two different questions:
Did something happen?
and:
Do we trust this measurement?
Those are not the same thing.
Making camera setup recoverable
Pose tracking also depends heavily on how the patient positions the camera.
We built a setup checker to decide whether the current view was usable for the exercise.
One issue was that early bad frames could influence the setup state for too long.
Imagine this:
bad setup
bad setup
bad setup
bad setup
patient fixes the camera
good
good
good
If the checker effectively remembers the entire history, the patient can remain stuck even after fixing the problem.
We changed the logic to focus on a recent rolling window instead.
That way, old bad setup frames do not permanently punish the patient.
We also kept checking the environment after the patient reached the Ready state.
Otherwise this could happen:
Patient is visible
โ Ready
Patient leaves the frame
โ still Ready
Instead, if the setup becomes unreliable again, Ready can be revoked.
We also added a timeout.
If the pose stream is running but the system cannot establish a reliable camera view after around 25 seconds, it stops waiting forever and shows the patient a retry option.
That turned camera setup from a one-shot gate into something that could actually recover from failure.
The patient needs to know what the tracker expects
Another bug was less about computer vision and more about product consistency.
An exercise level could require something like:
12 reps
70ยฐ movement target
2-second hold
The tracker knew about the hold.
The patient did not.
The UI originally might only say:
Raise your arm out to the side, then lower it slowly.
From the patient's point of view, they could perform the movement correctly according to the instruction and still fail the tracker.
We fixed this by deriving the patient-facing prescription from the same exercise-level configuration used by the tracking logic.
So the interface could show:
12 reps ยท Move through 70ยฐ at the shoulder ยท Hold for 2 seconds
and progress could show:
0 / 12 valid reps
If a level has no hold requirement, the UI does not invent one.
This gave us one source of truth for both tracking behavior and patient instructions.
The LLM never sees the raw video
One architectural decision I liked about Mendly was separating computer vision from LLM reasoning.
We do not need Gemini to inspect the patient's camera feed.
The CV layer can transform movement into structured information first.
Conceptually, the result can look something like:
{
"exercise": "arm_flexion",
"attempts": 9,
"valid_repetitions": 7,
"range_of_motion": 83,
"tracking_quality": "good"
}
The exact schema can change.
The important part is the boundary.
The computer vision layer answers:
What happened?
The AI planning layer answers:
How should that performance history influence a future exercise set?
The AI does not choose the next exercise after every rep
This is an important detail about Mendly's architecture.
The patient's active session follows an exercise set that has already been reviewed and approved.
Gemini is not sitting inside the rep counter deciding what exercise the patient should perform next.
The flow is closer to:
Practitioner defines boundaries
โ
AI drafts a future exercise set
โ
Deterministic guardrails validate it
โ
Practitioner reviews / edits / approves
โ
Patient performs the approved set
โ
Computer vision measures performance
โ
Validated results are stored
โ
Practitioner later requests another draft
โ
Backboard / Gemini use the accumulated history
That distinction was important to us because Mendly is not trying to make the LLM an autonomous therapist.
The model helps with planning.
The practitioner stays in control.
Where Backboard and Gemini fit
Our stack included:
- Next.js for the application
- MediaPipe for pose tracking
- Backboard for the AI orchestration layer and memory
- Gemini for adaptive planning
- ElevenLabs for voice-based accessibility
- TigerData for application data
- DigitalOcean for deployment and infrastructure
For the adaptive planning side, Mendly sends a planning request through Backboard, which routes it to a Gemini model.
Backboard also gives us patient-specific planning continuity.
But memory is not treated as the source of truth.
When a new draft is requested, Mendly still sends the current practitioner constraints and recent performance information.
We also kept a deterministic fallback path.
During testing, we actually hit a case where our Backboard API key worked, but LLM chat was unavailable because of account credit.
Instead of breaking the entire practitioner workflow, Mendly could fall back to a local rules-based proposal.
That experience reinforced a principle I want to keep using in future AI projects:
AI should improve a workflow, not become a single point of failure.
Human oversight matters
Because Mendly deals with rehabilitation, we did not want the model to have unlimited freedom.
The practitioner can define constraints such as:
- allowed exercises
- maximum exercise levels
- whether standing exercises are allowed
- affected side
- contraindications
- maximum exercises
- motor time limits
- rehabilitation goals
- practitioner notes
The model drafts within those constraints.
Then deterministic guardrails check the proposal again.
Then the practitioner makes the final decision.
The simplest way I think about it is:
Practitioner defines the box
โ
AI proposes inside the box
โ
Code checks the box
โ
Practitioner approves the result
That architecture was much more interesting to me than simply adding a chatbot to the application.
Safety support, not a "fall detector"
We also built a prototype safety-support layer around motor exercises.
The system can surface signals related to things like a rapid downward movement followed by a low body position, remaining low for an extended period, or disappearing from tracking for a long time.
Those events can create alerts for the practitioner.
But this is something I would describe carefully.
Mendly is not a clinically validated fall detector.
It is a prototype safety-support mechanism intended to surface potentially concerning situations to a human.
That distinction matters, especially in a healthcare-oriented project.
What changed how I think about computer vision
Before Mendly, it was easy for me to think of computer vision as:
input
โ
model
โ
prediction
After spending the hackathon debugging the movement pipeline, I started thinking about the whole system instead:
input
โ
prediction
โ
confidence
โ
validation
โ
temporal reasoning
โ
structured data
โ
application behavior
The model is only one component.
The engineering around the model determines whether its output is actually useful.
A system that sometimes says:
"I don't know"
can be much safer and more useful than one that always produces a number.
That was probably the biggest lesson I took away from working on Mendly.
What we built
By the end of HopHacks, Mendly had become a working prototype connecting:
Practitioner-defined constraints
โ
AI-assisted planning
โ
Practitioner approval
โ
Patient exercise session
โ
Camera-based movement measurement
โ
Validated performance history
โ
Future adaptive planning
Our team finished as a Top 11 finalist out of 87 projects and won Best Use of DigitalOcean.
But the result I care about most is that I left the hackathon thinking differently about AI systems.
The interesting part was not simply using MediaPipe or Gemini.
It was connecting the physical world to an intelligent system without pretending that noisy observations were perfect.
For me, the core problem became:
How do you turn imperfect observations into reliable information that an intelligent system can reason over?
That problem applies far beyond rehabilitation.
And it is probably the part of Mendly I enjoyed building the most.
Final Thoughts
Mendly started as a hackathon project, but building its computer vision pipeline changed how I think about AI engineering.
It is easy to focus on making models smarter.
But when AI interacts with real-world data, the quality of the reasoning depends heavily on the quality of the information we give it.
Sometimes the most important engineering decision is not:
"What should the model predict?"
but:
"Do we actually know enough to make a prediction at all?"
That was the lesson behind a lot of the fixes we made during HopHacks.
And it is one I plan to carry into the next system I build.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.