DEV Community

Cover image for I Built My Friend a Mock Interviewer That Read His Rust Code
Yash Kumar Saini
Yash Kumar Saini Subscriber

Posted on

I Built My Friend a Mock Interviewer That Read His Rust Code

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

grill is a mock technical interviewer that has read your code. It runs in a terminal, on an open model, on your own laptop.

I built it for my friend Akash. He has spent the last year writing Rust, ships TypeScript on the side, and is looking for a full-time role. The rounds he dreads are the Rust and JavaScript deep dives.

Most interview prep tools ask the same textbook questions: what is a closure, explain ownership. Real senior interviews do not go that way. They go "walk me through your project", and then the interviewer points at one line you hoped nobody would notice and asks why.

You cannot practise that with a question bank, because the questions depend on your code. So grill does three things:

  1. Reads his repositories and finds the lines worth a question: unsafe blocks, locks, lifetimes, spawned tasks, useEffect, forEach(async ...), any, empty catch.
  2. Interviews him on them. One question about that exact code, one follow-up that presses on the weakest part of his answer, then a grade from 0 to 4 with what he missed.
  3. Remembers what he got wrong. The next session draws more from his worst topics.

Here is the scanner on four of Akash's public repositories. This is real output:

Terminal output of grill ingest finding 116 interview-worthy spots across four repositories

116 spots, and the top one is in miccli, his voice dictation CLI: an Arc<Mutex<...>> shared with a real-time audio callback. That is exactly the line a Rust interviewer would circle.

Demo

[ADD: link to your terminal recording]

A session looks like this. The code and location are real scanner output; paste your own question, answer and grade from a real run in place of the placeholder lines:

── 1 of 5 ──
AkashJana18/miccli src/audio.rs:62  shared-state
62 │     pub fn start_capture(&self) -> Result<AudioStream> {
63 │         let (tx, rx) = mpsc::channel();
   │         ...
72 │         let sink = Arc::new(Mutex::new(PendingResampler::new(
   │         ...

[ADD: the question the model asked]
  > [ADD: the answer]

[ADD: the follow-up]
  > [ADD: the answer]

[ADD: the grade line, the "missed" lines and the "better" line]
Enter fullscreen mode Exit fullscreen mode

What Akash said

[ADD: Akash's reaction, in his words]

Code

grill: a mock interviewer that has read your code

grill

A mock interviewer that has read your code. It runs on your laptop, on an open model, with no network.

Most interview prep asks textbook questions. Real senior interviews ask "walk me through your project", then press on the one line you hoped nobody would notice. grill clones your repositories, finds those lines, and asks about them.

Built for Akash, who writes Rust and TypeScript and is interviewing for full-time roles.

I built my friend an interviewer that read his code

── 1 of 5 ──
AkashJana18/miccli src/audio.rs:62  shared-state
62 │     pub fn start_capture(&self) -> Result<AudioStream> {
63 │         let (tx, rx) = mpsc::channel();
   │         ...
72 │         let sink = Arc::new(Mutex::new(PendingResampler::new(
   │         ...

Why is `sink` behind a Mutex, and what does the audio callback do if that lock is contended?
  >

The code and location above are real scanner output. The question is an example of the kind a model asks; yours will differ.

Demo

grill interviewing a real repository: ingest, spots, a two-round interview, the weak-spot map

Recorded…




About 1,300 lines of TypeScript including tests, three runtime dependencies, MIT licensed.

ollama pull gemma3:4b

git clone https://github.com/yashksaini-coder/grill && cd grill
npm install && npm run build && npm link

grill ingest AkashJana18/miccli AkashJana18/solana-consensus-lab
grill start                    # 5 questions
grill start -n 3 --lang rust   # Rust only
grill report                   # the weak-spot map
Enter fullscreen mode Exit fullscreen mode

How I Built It

Architecture of grill: ingest, interview and remember stages, all running locally on Ollama, Gemma and Mastra

The open pieces:

  • Gemma, Google's open-weight model, served by Ollama on localhost.
  • Mastra, the open-source TypeScript agent framework, for the two agents and the tool.

Design rule: the model gets one small job at a time

A 4B model running on a laptop is not going to plan an interview, hold a rubric in its head, and remember what it asked five turns ago. So I did not ask it to. The order of steps lives in ordinary TypeScript. The model is only ever asked to do one of three things: write a question, write a follow-up, or grade an exchange.

One interview round: scanner shows a snippet, interviewer asks, candidate answers, interviewer follows up, grader scores, store saves

Every one of those calls is handed the real snippet again. The model never has to recall code, so it has less room to invent it.

Stage 1: finding the lines worth asking about

The scanner is a table of 38 patterns, 17 for Rust and 21 for JavaScript and TypeScript. Each one carries a topic, a weight, and the angle a good interviewer would take:

rust("lock", "concurrency", /\.(lock|read|write)\(\)\s*(\.await|\.unwrap\(\)|\?)/, 4,
  "how long the guard lives, whether it is held across an await, and what a poisoned or contended lock does"),

js("async-in-loop", "async", /\.(forEach|map|filter|reduce)\(\s*async\b/, 5,
  "what actually awaits here, the order of completion, and how errors surface"),
Enter fullscreen mode Exit fullscreen mode

A regex hit on one line is not a question. So for every hit, the scanner walks up to the enclosing function, matches braces to find its end, and frames the whole thing. Weights are summed per function, overlapping frames are dropped in favour of the strongest, and each repository contributes its best 40 spots, round-robin across topics so a codebase full of .unwrap() does not turn the whole interview into error handling.

This is regex and brace counting, not a parser. It finds interesting code. It does not understand it. Understanding is the model's job.

Stage 2: two Mastra agents on Gemma

Mastra talks to Ollama through its OpenAI-compatible endpoint, so there is no provider SDK and no API key:

const model = { providerId: "local", modelId: config.modelId, url: config.url, apiKey: "local" };

this.interviewer = new Agent({ id: "interviewer", name: "interviewer", instructions: INTERVIEWER, model });
this.grader = new Agent({ id: "grader", name: "grader", instructions: GRADER, model });
Enter fullscreen mode Exit fullscreen mode

The interviewer is told to ask exactly one question, to name an identifier from the snippet so the candidate knows it read the code, and never to answer or hint.

The grader is a separate agent with a separate prompt. It did not write the question, so it has no stake in the answer being good. It returns JSON:

{"score": 2, "verdict": "one sentence", "missed": ["specific point"], "better": "the answer you wanted"}
Enter fullscreen mode Exit fullscreen mode

Small models wrap JSON in prose and code fences. Instead of trusting the reply, grill pulls out the first balanced {...}, validates it with zod, and if that fails, shows the model its own bad reply and asks once more.

There is also an opt-in read_source tool (grill start --tools). It lets the interviewer read beyond the snippet, for example to see a caller or a type definition. It is confined to the repository: a path that resolves outside the checkout is refused. It needs a model with tool calling, so it is off by default.

Stage 3: the weak-spot map

Every round is saved as JSON under ~/.grill/sessions. The next session picks spots with a weighted draw:

export function priority(spot: Hotspot, seen: Set<string>, stats: TopicStat[]): number {
  const stat = stats.find((s) => s.topic === spot.topic);
  const weakness = stat ? (4 - stat.mean) / 4 : 0.5;
  const repeat = seen.has(spot.id) ? 0.15 : 1;
  // Squared, so the first session opens on the richest code, not on filler.
  return spot.score ** 2 * (0.5 + weakness) * repeat;
}
Enter fullscreen mode Exit fullscreen mode

Score a 1 on lifetimes and lifetimes come back. Score a 4 on unsafe and it fades. Spots already asked are rare but not banned, because being asked the same thing twice is how you find out whether you learned it.

Testing an LLM app without an LLM

I did not want the test suite to need a GPU. The tests start a fake OpenAI-compatible server on a random local port and point the real Mastra agents at it. That covers a full round, the JSON retry, :skip and :quit, and the tool call including the path-escape refusal. 18 tests, a few seconds, no model.

Why Does Open Innovation Matter?

For this project it is not a preference. A closed API would have made it a worse tool for Akash in four concrete ways.

His fumbled answers stay his. Interview practice only works if you are willing to be wrong. grill stores every weak answer he gives, by design. That file sits in ~/.grill on his disk. With a hosted API, the same record of everything he does not know would be on someone else's server.

Private code works. grill ingest takes a local path. Work code under NDA, the repository he cannot push anywhere, the side project he is not ready to show: all of it can be interviewed on without leaving the machine. Once the model is pulled, a local-path session needs no network at all.

Practice is free, so he can do a lot of it. Three model calls per question, five questions per session, as many sessions as it takes. On a metered API that is a bill that grows with exactly the behaviour I want to encourage.

He can change the interviewer. The rubric is a string in src/interview/mastra-brain.ts. If he is preparing for a company that cares about system design more than language trivia, he edits the prompt. If he gets a better laptop, GRILL_MODEL swaps in a bigger model in one environment variable. If Ollama is not his thing, GRILL_URL points at llama.cpp or LM Studio. Nothing in the tool is tied to one vendor.

Open weights made it private and free. An open framework made it changeable. He needs all three.

Limits, honestly

  • The scanner cannot tell a deliberate unwrap() from a careless one. It surfaces candidates; some are dull.
  • Question quality is the model's. A larger local model asks sharper questions.
  • The grade is a study aid, not a verdict. Read the "missed" list, not the number.

Prize Categories

  • Best Use of Gemma. Gemma runs locally through Ollama and does all three jobs: asking, pressing, and grading. The whole design, with fixed orchestration and one small task per call, exists to make a small open-weight model enough.
  • Best Use of Mastra. Two Mastra agents over an open model, plus a sandboxed read_source tool, wired to a local OpenAI-compatible endpoint with no provider SDK.

Top comments (1)

Collapse
 
rohan_sharma profile image
Rohan Sharma • • Edited

It's looks so good, Yash.

btw, I found some placeholders:

github link is also broken: github.com/yashksaini-coder/grill/..., it should be github.com/yashksaini-coder/grill/