DEV Community

Alex Georgiev
Alex Georgiev

Posted on

Built anything cool with AI lately?

From fitness apps to custom programming languages

I train a lot, and with AI's help I built a small app that pulls my Garmin and Strava data and helps me structure my training around it.

Made me curious what else people are building. What's something you've put together recently with AI's help, big or small?

Top comments (29)

Collapse
 
chervet profile image
Cédric Hervet, Ph.D. •

I used an AI coding agent to help build the piece that lets other AI agents reliably call our route optimization API. Kind of meta watching my coding agent debug why a different agent kept picking the wrong constraint model. Turns out the hard part wasn't the API call, it was the agent misunderstanding which business rule applied. Ended up spending more time making the docs "agent-readable" than on the API logic itself.

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

That's a great example of the gap being in shared understanding, not the interface. Makes sense the docs ended up being the real work, an API can be technically correct and still get misread by an agent that doesn't have the business context a human would infer. Which model was doing the debugging, and did it figure out the fix itself or just surface the disagreement?

Collapse
 
chervet profile image
Cédric Hervet, Ph.D. •

We did it with Claude Opus, and we intend to generalize the process to other models in order to provide some efficiency scores per model to our users. The process looks like this: we have an agent trying to build a correct integration based on some data, our docs, and a short brief. Then another agents reviews the optimization model against a target model built by our experts for the toy problem, and suggests improvements in the documentation to avoid pitfalls or guide the agent more. Then we start over with the updated doc to assess the improvement. Simple, but cool to watch. Of course we keep humans in the reviewing loop to make sure we do not degrade our docs.

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

Nice, that's a solid loop. One thing to watch: docs tuned to fix Opus's blind spots might not transfer cleanly to other models. Their pitfalls don't always overlap. Might be worth tracking whether the fixes are model-specific or general doc gaps.

Thread Thread
 
chervet profile image
Cédric Hervet, Ph.D. •

Yeah definitely, that's why we want to expand it to other models asap. Claude was a natural choice since our surveys shows that our users are massively relying on it (more than 75% of them are). Surprisingly enough, the level of non determinism in Claude responses is quite high. At each iteration, we launch 5 to 10 runs in parallel and it's crazy how innovative the same Claude model can be in front of the same initial prompt. So basically, it's not about the documentation, it's about multiplying touchpoints (api reference, sdk docstrings, modeling guides, nudging messages in API responses, etc.) to make sure that whatever the path chosen by the agent, it will end up reading the right piece of documentation to solve the situation.

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

Makes sense, redundancy beats a single perfect doc if you can't predict which path the agent takes. Curious how you settle on a score when 5-10 runs diverge that much, best-of, median, or something else?

Thread Thread
 
chervet profile image
Cédric Hervet, Ph.D. •

Well, we use the mean to measure the global efficiency of agents, but what we really look at is how many runs are completely failing at delivering "something useful". We're still iterating on that notion. What I can say is that "something useful" does not mean "perfect", but rather a "good basis to iterate on". Indeed, even when humans are building models, it takes a few shot before getting it right, so it's only natural that agents may require a few turns before landing on a model that could work in real life.

Actually, having no iteration at all would be suspicious, as operational team always come up with something new when they meet the computed consequences of their statements. That's okay, that's part of the process, that's how you go from informal "best" practices to a structured, well defined process.

So what matters mostly is that on the first run, the agents does not go in a totally wrong direction, and that most of its choices will deliver a good model to reflect on, so future iterations are focused on fine tuning it rather than debugging or worse, restart from scratch.

Therefore, we have two KPIs we look at : % of totally failed runs, and average score. The first one is the most important to work on.

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

That tracks, a low failure rate matters more than a high mean if a few good runs are just masking one disaster. What counts as a totally failed run, no working code at all, or something that compiles but misses the brief entirely?

Thread Thread
 
chervet profile image
Cédric Hervet, Ph.D. •

Getting a working code out of an agent is not the hard part. Especially in route optimization, it's very common to get a 200 in return of your optimization request, but still getting something that does not make sense operationnally. Request is valid, but the problem formulated by the agent makes no sense. That's the kind of silent failure we want to avoid

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

That's the nasty one, a 200 with a semantically broken formulation won't show up in any error log. How does the reviewing agent catch that, domain specific checks per problem type, or something more generic?

Thread Thread
 
chervet profile image
Cédric Hervet, Ph.D. •

Basically, we provide the reviewing agent a "perfect" model built by our own team, with a markdown giving all the details of each choice we made. So the reviewing agent compares both and provides feedback based on that. Pretty much a supervised learning approach where we expose our homemade "ground truth"

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

That's basically supervised learning with your team as the labelers, makes sense given how hard "correct" would be to define otherwise. Really enjoyed this thread, thanks for walking through the whole pipeline. Looking forward to catching up again soon!

Thread Thread
 
chervet profile image
Cédric Hervet, Ph.D. •

Yup, occam's razor, we've been doing this for more than 10 years, on countless projects with different complexity levels. So let's be the training set ourselves :)
Enjoyed it too!

Collapse
 
botsailorofficial profile image
BotSailor •

I really enjoyed reading this. I’m still learning and exploring what’s possible with AI, so seeing the different ways people are experimenting with it is genuinely inspiring. I think the best part is that you don’t need to build something huge or perfect to learn something valuable. Even a small idea can teach you a lot when you actually try it. Thanks for sharing this, it definitely gives me a few ideas to explore myself.

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

Glad you enjoyed reading it :) . Small ideas are honestly the best way to learn, what you figure out along the way matters more than whether the thing itself is impressive.

Collapse
 
unitbuilds profile image
UnitBuilds •

My fiance wanted a fitness app, so I built 1, decided to add a bit more to it though... So it has a digital pet, AI trainer, mental health section, cycle tracker, smartwatch integration and a few more bells and whistles.

I'm also busy refining my NDA based OS, a SASOS built on a custom language, that runs a ring-0 LLM harness in the kernel, so it can manage it's own hardware

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

That sounds really cool! I hope she enjoys it!

Mine has AI trainer, I track basically everything when it comes to data, I compare effort, I analyze data, gives health data like fatigue and etc.

Your side project also sounds nice! :)

Collapse
 
unitbuilds profile image
UnitBuilds •

My idea with the pets was that it's essentially your 'goal', what animal best embodies how you want to be. Eg. if you're a bit shy and want to boost self-confidence, choose a dragon, so through aligning with it, you learn to be more self-confident. If you're anxious and want to learn how to roll with it more, choose the panda, so through alignment, you learn to be calmer and more carefree.

I also made it so event planning, you can pick your dates ahead of time, set your expectations, then it'll adjust your plan leading up to it, so going out for drinks on the weekend isnt a guilty pleasure you need to pay for afterwards, instead you've prepared for it, so it's a reward for putting in the effort ahead of time.

The AI trainer also judges your state and history, so you can ask it questions, eg. if you're feeling meh, ask the trainer if you can skip your workout, or if you really want pizza, whether you're allowed it. It wont always say no and it wont always say yes, it tries to comply with your lifestyle, unless it detects you're slipping and need to be realigned. Alot of thought went into it's systems so it actually acts like a proper trainer, not just a yes-man or bluntly enforces rules.

Also added meal planning and groceries, so it can plan out a grocery list for you and learns from your receipts what's cheap, what's expensive, what's seasonal, etc. So you can get nice meals, without breaking the bank. It tracks per meal what your usage is, so it knows you have 2 tomatoes left. That obviously wont work if you dont eat alone, so I added family tracking too, either full tracking, or as little as just usage based RL from when you notify it you're out of something, it knows your meal was 2 tomatoes, if you ran out of your 5 tomatoes, it means your family ate 3, etc.

The mental health section was also quite important to me, it lets your pet pay attention to your biometrics and if you deal with eg. anger issues, or anxiety, it can detect your biometrics and interact over your smartwatch to just help you through the situation.

I'm looking at maybe releasing it on the stores, but I hate subscriptions... And ads... So I wanna make it BYOK, else use my key and pay for the AI usage. That way it's lean and doesnt force anyone into paying a penny more than what they use.

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

Love the BYOK instinct, no ads no subscriptions is the dream. I went a similar route with mine, my friends get full access for free, and I opened it up publicly too with a free tier that's just limited features rather than a paywall. Keeps it lean like you said, and the people who really get value out of it tend to stick around anyway.

Collapse
 
edmundsparrow profile image
Ekong Ikpe •

An SVG editor. Basically reads svg and edits OTG. Not export but save back to SVG. Still exploring

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

That sounds nice, I would guess this was giving you a hard time doing it manually before?

Collapse
 
edmundsparrow profile image
Ekong Ikpe •

More about exploration. SVG as a CAD app tool not an export.

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

ah, got it now. I've also created a lot of tool for exploration, most of them tuned out to be absolute surprise to me, wondering what my inital prompt was :D :D

Collapse
 
skymonder-alt profile image
SkyMonder-alt •

I've been using Claude to build a Russian-syntax programming language
(SkyForge). The interesting part for me wasn't writing the interpreter —
it was building a transpiler that converts my AST into Python AST so
compile() can turn it into native bytecode.

fib(32) went from ~60 seconds (tree-walking interpreter) to ~0.4
seconds (transpiled). That's within 6% of plain CPython.

The AI was genuinely useful for two things: boilerplate (writing 30+
ast.* node types by hand is tedious) and catching weird edge cases
I'd miss — like Python 3.12+ refusing to compile an AST where a child
node's lineno is less than the parent's.

Hardest part: AI tends to invent keywords that don't exist in my
language. Had to write a strict context file of "only these constructs
exist" and it still happens occasionally.

Collapse
 
alexgeorgiev17 profile image
Alex Georgiev •

That's a nice trick, going through Python's own AST and compile() instead of writing your own bytecode compiler basically gets you CPython's actual VM for free. 0.4s for fib(32) within 6% of native CPython is a great result for the code you didn't have to write yourself.

The lineno ordering constraint is a fun one to hit blind, that's a genuinely obscure CPython internal to stumble into from a hobby language.

On the invented-keywords problem: have you tried feeding the model a formal grammar file (even a stripped-down EBNF) instead of, or alongside, the prose "only these constructs exist" context? A grammar tends to anchor these things harder than a list of allowed words, since it also constrains structure, not just vocabulary.

Collapse
 
skymonder-alt profile image
SkyMonder-alt •

Fair point — I haven't tried EBNF yet, which is honestly an oversight
given that I wrote the parser by hand and could have derived the
grammar from it.

What I use instead: a short list of canonical examples, one per
construct. Instead of "these keywords exist", the context file has
a small snippet for функция, если, для, попробовать, etc.
That anchors better than prose because the model copies structure,
not just vocabulary — you can't fake a valid если block without
also writing иначе in the right place.

But you're right that a grammar would be stronger. The thing I keep
hitting is structural drift, not word invention: the model gets the
keywords right but nests a block one level too deep or forgets a
closing brace. A grammar constrains that; examples don't.

Did you have a specific approach for feeding EBNF to a model in
practice, or is it just "paste the grammar in the system prompt"?

Thread Thread
 
alexgeorgiev17 profile image
Alex Georgiev •

Just pasting the grammar into the system prompt is closer to a suggestion than a constraint, the model isn't checking its own output against it token by token, so you get exactly the failure mode you're describing: it "knows" the grammar in the sense of pattern-matching surface structure, but nothing stops it from drifting a level deep on nesting or dropping a closing brace, since there's no verification step.

Two things actually close that gap, and they're different tools for different setups:

If you have access to constrained decoding (llama.cpp's GBNF grammars, or something like the outlines/guidance libraries on top of a local model), you can restrict token sampling so only grammar-valid continuations are ever possible. That's a hard guarantee, not a hope, structural drift becomes literally unsamplable.

If you're on a hosted API without that option, the more practical pattern is generate-then-repair: run the output through your own real parser (the one you already wrote by hand), and when it fails, feed the model the exact error and location rather than asking it to "try again." A real parser catching an unclosed иначе at line 12 is a much sharper signal than the model silently re-guessing.

Given you already have a hand-written parser, the second option is probably the lower-effort win: you already have the ground truth, you're just not closing the loop with it yet.

Collapse
 
syke99 profile image
Quinn Millican •

I just released the beta of Natyv, a desktop development framework that doesn’t rely on a WebView or DOM after about a month of using Claude Code to do the grunt work while I spent my time architecting it all.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.