Steadywag: A Care Agent That Knows What His Paperwork Gets Wrong
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content A medication was on the written list. The dog never got it. A specialist told one caregiver out loud to skip it. Nobody u
This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
A medication was on the written list. The dog never got it.
A specialist told one caregiver out loud to skip it. Nobody updated the paper. A second caregiver covering his meds would reasonably think it was current, because every document said so.
That gap is the whole project.
What I Built
Steadywag is a care companion for a dog with a chronic illness, built on one real case: Theo, an 8-year-old Shih Tzu mix with copper storage hepatopathy, recurrent pancreatitis and high triglycerides. His care lives in about 180 pages of PDFs, emails and things said out loud in exam rooms.
I pulled it into one structured chart in Sanity, then built the app on top of it. It installs on a phone or a computer, works offline for the pages you have opened, and follows light or dark mode.
- Ask. Pick a task: What does he need today? Something changed. Can he eat this? I'm watching Theo (an instant sitter brief you can print). Prep my vet visit (with your own check-ins). Or ask anything. While the agent works, you watch it: "Reading his medication list," "Reading his family routine." When it finishes, source chips link to the pages it used, and you can keep the conversation going with follow-up questions.
- Today. The day as a timeline, built from data with no agent call. It follows the clock: "Next up: dinner at 8 PM," then "Later today," with finished steps collapsed. His vet's written instruction comes first. The family's own routine fills in a time only where the vet gave none, and every time is labeled. The medication that is listed but not given sits at the top, folded to one line so it does not take over the screen.
- Check-in. A 30-second daily log of the things his vet asks about at every visit: eating, drinking compared with usual, energy, stool (picked from pictures), vomiting, meds, activity, and a body check for bruising or a yellow tint in the eyes, ear flaps or gums. Before a visit it counts them across days: "drinking more than usual on 2 of 3 days," counting only the days a question was answered.
- Appointments. Vet, grooming and dental visits, with a month calendar and an "Add to my calendar" file that includes a reminder the day before.
- History. "Theo in 60 seconds," one chart with ALT, medication periods and flares on a shared axis, five chapters and four patterns. Each claim points at the records it rests on. A "What changed" log and his DNA test report sit alongside.
- Your dog. A pre-diagnosis diet questionnaire, so families can give researchers consistent dietary copper data. It starts blank. Answers never leave the visitor's browser.
What the agent will and won't do
It never diagnoses, never gives or converts a dose, and never tells you to start or stop anything. It reads. It cannot write.
- Verdict first. For "can he have X," it opens with one verdict: Fits his plan's rules, Probably not, No, or Can't tell. Then it shows the checks against his own plan: copper compared with the foods already approved, sodium, fat, calories against the treat limit, plain preparation, and his avoid list. It never says "safe," always says when his vet has not confirmed it, and ends with a question to ask his vet.
- Two narrow reaches outside his chart. For a food his plan does not cover, it looks up copper, sodium and fat in USDA FoodData Central, labeled "USDA food data, not from his vet." For general questions it can run a web search, at most twice per answer, limited to a short allowlist of veterinary and nutrition sites. Anything from the web is labeled "From the web" with a link, and nothing about medication or doses is taken from it. His vet's plan always wins.
- Asks for the ingredients. If someone asks about a store-bought product, it does not pretend to search for the label. It asks for the ingredient list and checks that.
- Stays on topic. A small, fast model checks that a free question is about Theo, dogs or the app. Anything else, like math, a poem or "ignore your instructions," is declined in one sentence in about half a second.
Demo
- Live site: https://steadywag.com (the original https://steadywag.vercel.app also works)
- Video walkthrough: https://www.youtube.com/watch?v=LfrX6UZ_XJU
Try these on the Ask tab:
- What does he need today? It opens with the medication that is on the list but never given.
- I'm watching Theo. It opens an instant sitter brief that starts with the medication the paperwork gets wrong, and ends with a reminder to add phone numbers.
- Are his freeze-dried treats allowed? His records disagree. The answer shows both sides and refuses to pick one.
- Can he have mango? His plan does not cover it, so you can watch it look up the USDA numbers and compare them with his approved foods.
On every page except Ask, a floating chat button opens the same agent in a window, so you can ask while you read.
Code
narlynars07
/
steadywag
A care companion for a dog with a chronic illness. An agent on Sanity Context that tells you what the records don't say. DEV Sanity Challenge, Path One.
Steadywag
A care companion for a dog with a chronic illness. It reads his records and tells you what the paperwork doesn't say.
Built on one real case: Theo, an 8-year-old Shih Tzu mix with copper storage hepatopathy, recurrent pancreatitis and high triglycerides. His care lived in about 180 pages of reports, emails and things said out loud in exam rooms. Steadywag turns that into one structured, sourced record in Sanity, then puts an agent, a daily plan, check-ins and appointments on top of it. Built for the DEV Sanity Challenge, Path One: an agent that queries real content.
The idea in one example. A medication was on the written list, but the dog never got it: a specialist told one caregiver out loud to skip it, and nobody updated the paper. The structure knows (medication.status = "listed-not-given" plus a recordGap). The PDFs don't. The agent, theβ¦
Next.js 16 and the Vercel AI SDK on the web side. A Sanity Studio and a 22-type schema in studio/. The README explains the architecture and how to run it, and evals/ holds the audit questions described below, with a runner.
How I Used Sanity
The structure does the work. Here is what a keyword search over the original PDFs could not tell you:
-
medication.status = "listed-not-given"plus arecordGapof kindverbal-instruction. The PDFs show the drug as active. The fact that it was never given exists only because the family told me. -
sourceNote.confidenceon every fact from a record:confirmed,single-source, orconflicting. A conflict is a queryable state, not a sentence buried in a note. Cerenia is one: his medication list says Monday, Wednesday, Friday, and the written instruction in the same report says every 24 hours as needed. Both stay visible. -
recordGapdocuments. Absence cannot be searched for. It has to be modeled. The chart holds 9 open gaps, like the lab result a specialist said was pending that never arrived. -
dietHistoryEntry.origin:family-recallormedical-record. The agent has to say which one it is quoting. -
medication.writtenInstructionnext to a separatecareRoutinetype. The vet's words and the family's routine are different sources. The agent labels each one: Vet records, Family routine, Family recall, Family observations, Your check-ins, From the web. When the vet gave a time, it uses it. The routine never overrides it. -
familyNotedocuments: what one family has noticed, like the first signs of a flare. They say whether his vet directed them, and the agent is told never to turn them into advice. Medication timing and doses are kept out. -
historyChapterandhistoryPatterndocuments, each holding references to the visits, labs and medications it rests on. The History page and the agent read the same documents. -
dnaReport: his consumer DNA test exactly as the report states it. Putting it next to the guidance gave the agent a cross-source answer: the report has no copper-related test, and none of his breeds is among those the copper guidance names. Both facts are published and labeled. Kit numbers and personal links are never stored.
The dataset is public: 516 documents across 22 types, including 288 lab results, 26 visits and 26 weights. ALT went from 2,431 in December 2023 to 59 in August 2026, and you can chart that from the raw data.
Two Sanity Context endpoints. The agent connects to both:
- A chart endpoint in GROQ mode over the dataset. Tools:
groq_query,schema_explorer,array_field_reader. - A knowledge endpoint backed by a Knowledge Base. I pointed it at the dataset with one GROQ query,
*[_type == "guidance" || _type == "dietRule"], which matches 26 documents: the 10 citedguidanceentries and the 16dietRuledocuments. Sanity builds entries from them and flags conflicts in an Issues view. Tools:knowledge_base_search,knowledge_base_read. What the agent does with what it retrieves: it quotes the published guidance next to his own plan, names the source, and says when the evidence is not about his breed.
The Knowledge Base earned its place, and it also failed in two instructive ways. It gave the agent published guidance to cite next to his own plan, like the ACVIM consensus statement on copper-associated hepatitis. But it also invented a "400-600 borderline" row in a liver-copper threshold table that no source gives, and the agent repeated it. And its conflict check read a date written "3/10" as 3 March. I fixed the first with a standing Knowledge Base instruction and the second by spelling out months in the source text.
The agent is Claude Sonnet 5.5 with up to 12 tool steps and read-only access. Each task adds its own instructions on top of the base rules. For "Something changed," warning signs come first, and past flares are framed as "During past flares, his records show..." and never as instructions for now. It looks things up first and cites what it used. If the Context endpoints are unreachable it falls back to querying the same public dataset directly, but never silently: the answer shows a "Direct query, not Sanity Context" notice.
Visitor data stays out of the dataset. Check-ins and appointments live in the browser. That is deliberate: they are private health details, and saving them on a server would need accounts, because without sign-in every visitor's entries would land in one shared pile. Each page says "Saved on this device only," and a backup file moves them to another device. When you ask for a visit summary, the browser sends the entries since his last visit in a validated, size-capped field. They go to the model marked as family-entered observations, inside delimiters, as data. Nothing is stored server-side, and the daily question limit keeps only a one-way hash of the visitor's address, never the address.
Where it broke, and how I tested it. I tested hard, and the questions are in the repo.
evals/ holds 24 audit questions, each with the reason it exists and a few checkable expectations, plus a runner (node evals/run.mjs). They are smoke checks, not a grade, so I read the answers too. Two rounds caught real problems:
Round one: 12 questions, 4 failed.
- Blueberries. The agent said his plan does not name them. His nutrition consult lists them as an approved treat. My food documents carried no source note, so the agent could not tell where the plan came from.
- What he ate before diagnosis. The agent could not see the answer. I had hard-coded those diets in the page instead of storing them in Sanity.
- Liver copper. The invented Knowledge Base row, above.
- Atopica. One answer claimed nothing was on file after 2024. The data was there. My query never asked for the confirmation date.
Round two, after the rebuild.
- Empty answers. "What does he need today?" and the sitter brief came back blank. The agent spent all 8 steps on lookups and never wrote. I raised the cap to 12 and made the final step text-only, so an answer always gets written.
- Cut-off answers. Long briefs stopped mid-sentence and lost their closing line. I raised the output limit and capped the sitter brief at about 350 words.
- The same filter trap, twice. The chart endpoint filters by document type. My new routine and history types were invisible to the agent until I added them, and so were the family notes. It said, correctly, "I couldn't retrieve his family routine." That honest failure is how I found it both times.
After that I found two more by using it, not by test: a store-bought product answer where the agent talked as if it could read a label, and web results that read as if they came from his vet. Those became the ingredient-list rule and the "From the web" labels.
The final run passes all of it through Sanity Context. No answer diagnosed, gave a dose, invented a time or named a clinic or person. Median answer time was about 18 seconds.
A record that stays true. A record is only useful if it changes when the facts change. If his vet removes ursodiol from the written list, the owner marks it stopped in the Studio, resolves its gap and its question, and adds an entry to the What changed log, with who said so. Within about a minute the home page callout, Today, the sitter brief, Visit prep and the agent's answers all stop flagging it, because every one of them reads the same record. Nothing on the website edits the record from the browser, so every change has a person and a source behind it. The steps are in docs/keeping-the-record-current.md.
Sanity Project Details
-
Project ID:
yahsq70q -
Dataset:
production(public read) -
Dataset query URL:
https://yahsq70q.api.sanity.io/v2026-10-01/data/query/production?query=count(*) - Studio: https://steadywag.sanity.studio
- Context endpoints: two (chart in GROQ mode, knowledge over a Knowledge Base of 26 documents)
-
Document types: 22, including
dnaReport,recordUpdate(the change log),careRoutine,familyNote,historySummary,historyChapterandhistoryPattern
What I would tell another builder
Model the sources, not just the facts. Once the vet's words and the family's routine were separate documents, the agent stopped blending them.
Keep the gaps. They are the product.
Embeddings are not enabled on this dataset. Retrieval from the Knowledge Base is keyword search, so it only works because the content behind it is structured and cited.
The guidance entries are summaries in my own words with a link to every original. A veterinarian has not reviewed them yet, and the site says so.
The records are de-identified: first name and birth year only, no clinic, clinician, owner or account details.
His medical record is shared: it lives in Sanity, so every device sees the same thing. His check-ins and appointments stay in your own browser on purpose. In a real product, sign-in and a private database would sync them. I kept the demo private instead of shared.
Credits. Nutrition numbers come from USDA FoodData Central. The guidance entries summarize published veterinary sources, each linked in the dataset, including the ACVIM consensus statement on copper-associated hepatitis. Built with Next.js, the Vercel AI SDK, Claude, Sanity Studio and Sanity Context. Steadywag tracks and prepares. It is not medical advice.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.