Rootstrap Get Started →

Insights

How Rootstrap Hires for AI Fluency, Not Just AI Use

How Rootstrap Hires for AI Fluency, Not Just AI Use

A while back, something happened that stuck with me. A candidate made it pretty far in our technical hiring process, passing stage after stage, and at some point we realized something: they barely used AI. What they used was mostly ChatGPT, but they had no real experience with code assistants or any other AI tooling. And at that stage of the process, given the level we expect today at Rootstrap, that wasn't enough anymore.

It wasn't that our process was broken. We had been working on this internally for a while already, we knew the developer role was changing fast and that whoever joined needed to be aligned with that from day one. But that specific case showed us something more precise: we weren't evaluating this dimension well enough. We needed to move as fast in how we measure as the tools we're measuring against are changing.

What we ended up doing was turning AI Fluency and a real product and business mindset into hard requirements for every engineer we hire, not soft preferences somewhere in the mix.

The problem wasn't "they don't use AI"

The first thing we had to understand was that the problem wasn't binary. It wasn't about separating candidates who use AI from candidates who don't. It was something more subtle: some people use AI and produce weak, ungrounded output, without being able to explain why they made the choices they made. And some people use it the way it should be used, with judgment, knowing when to trust it and when not to.

Our old process was built to check whether someone could write correct code. That goal is degraded today. Anyone can arrive at a result that "works" with AI's help. What we needed to measure was something else: technical judgment, the ability to stand behind your own decisions, and accountability, taking ownership of the outcome no matter who or what generated it.

Giving AI Fluency a real definition

One of the clearest things that came out of all this is that we needed to stop treating "AI Fluency" as a vague, feel-good phrase and actually define it. Not just "uses AI tools," but a real, observable skill with levels attached to it.

We looked at how other companies in the industry were approaching this, and a lot of that research pointed back to the same underlying framework: Anthropic's 4Ds of AI Fluency.

  • Delegation, knowing what to hand off to AI and what not to
  • Description, communicating a problem clearly enough for AI to actually help
  • Discernment, critically evaluating what comes back
  • Diligence, taking responsibility for what you ship, regardless of what generated it

We didn't adopt that framework as-is. We used it as a starting point and built our own version around it, one that reflects how we actually work with clients, in shifting stacks and domains, with a lot of autonomy and real business impact riding on the decisions we make. What came out of that is a leveled scale we now use across the hiring process:

  • Not using it, no real practice with AI tools yet, sometimes framed as a personal choice, but the pattern underneath is avoidance.
  • Basic, reaches for AI here and there, but takes what it gives back without much question, and can't really account for the parts they didn't write themselves.
  • Solid, already thinks in agentic workflows, not one-off prompts, catches when an output looks right but isn't, and steps in before it becomes a problem down the line.
  • Advanced, doesn't just get agents to work, builds the way of judging whether they're actually working, defines the guardrails upfront, and can explain a failure instead of just patching around it until it passes.

This is a snapshot, not a fixed bar. What counts as Solid or Advanced keeps moving as fast as the tools do, so we revisit these constantly instead of treating them as settled. And the lowest level isn't just "not ideal," it's disqualifying. That was a deliberate choice. This isn't a nice-to-have anymore, it's part of the baseline we hire against.

Why product judgment matters as much as AI skill

Something else became clear as we went through this: AI Fluency wasn't the only gap. As AI takes over more of the mechanical work, what starts to matter more is whether someone actually understands the product and the business behind what they're building, not just the technical problem in front of them.

We started paying much closer attention to what we internally call a product and business mindset. Does this person question a requirement when something doesn't add up, or just build exactly what's written down? Have they actually been part of conversations about why something is being built, or have they only ever received instructions and executed on them?

This isn't unrelated to the AI conversation, it's the same underlying shift. When writing code stops being the bottleneck, what differentiates someone is the judgment they bring to everything around the code: what to build, why, and for whom. So this became a second dimension we now evaluate deliberately, alongside the developer role was changing fast, instead of treating it as something that would just show up naturally somewhere in the process.

How we changed our hiring process

The spirit behind the changes was consistent across the board: stop measuring output, start measuring judgment.

We adjusted how we observe someone working through a technical exercise, giving them room to interact with AI the way they actually would on the job, instead of treating that as something to hide or work around. We also adjusted how we talk with candidates afterward, focusing less on whether the final result was "correct" and more on how they got there: what they questioned, what they accepted, what they'd do differently.

What we're actually looking for

What separates someone who "uses AI" from someone who's genuinely fluent isn't how many tools they mention in a conversation. It's whether they can describe a time AI got something wrong and how they caught it. It's whether they can tell you which tasks they delegate and which they don't, and explain why.

But there's something we've come to weigh even more heavily than a snapshot of where someone is today: the trajectory that got them there. Two candidates can look identical in an interview, same tools, same vocabulary, and still be in completely different places. One might have picked up their current workflow eight months ago and never touched it since. The other is actively experimenting, dropping things that didn't work, picking up new ones, and adjusting constantly. That second person is the one we want, even if today, on paper, they look the same as the first.

Today we see real cases, without going to extremes, where candidates don't move forward exactly for this reason. Not because the final result was wrong, but because there's no visible judgment in how they got there. They accept the first output without questioning it, or can't explain a decision beyond "that's just what the AI did." Looking at last month's technical challenges alone, 29% showed strong judgment from the very first technical stage, and 12% showed close to none of it, low enough that it's often where the process stops for them.

There's one idea that became central for us, and it connects directly to how we work in delivery: a spec-driven approach, where the person defines the contract before letting AI iterate. That means deciding what's critical to validate, what risks actually matter, and then judging whether what got generated truly meets that or just looks like it does. It's not about writing every line by hand. It's about knowing where to draw the line, when to push back, and when a result needs more work before it gets accepted. The same logic applies on the product side: knowing where the real requirement is, even when what's written down says something else.

What's still evolving in this framework

This isn't a redesign that closes and stays fixed. We know the tools are going to keep changing, and the process has to keep pace without falling behind, the way it did before.

There's still plenty we're actively refining, both on the AI side and the product side, adjusting what we look for as we learn more from every round of candidates, and building exercises that get even closer to what someone will actually face on a real project.

What we are sure about is the direction. We're not looking for people who know how to use one particular tool, or who can recite the right answer in a requirements meeting. We're looking for people who can hold onto their own judgment in a context where AI is doing more and more of the mechanical work, and where understanding the business behind the code matters as much as the code itself. That's the filter that matters to us, and it's the one we'll keep sharpening as all of this keeps changing, because we have no doubt that it will.

Why this matters beyond hiring

Clients sometimes ask, fairly, how they can know that whoever we staff on their project actually knows how to use AI well, and isn't just saying so. This whole redesign is really our answer to that. When a client brings us onto a project, they're trusting that whoever we staff actually knows how to work this way, not just talk about it. That's the whole point of raising the bar here: judgment is something we want to guarantee upfront, not something the client discovers mid-project. Our partnership with Anthropic to get the team certified is part of that same commitment, because this needs to be real on a client project, not just sound good in an interview.

That candidate from the beginning of this post wouldn't slip through unnoticed today. Every single engineer we evaluate now gets mapped against a clear AI Fluency level and a product and business mindset level, no exceptions. It's not a finished system, and we don't expect it to ever really be finished, but it means we're not caught off guard at the finish line anymore.

← All insights