The Roles AI Teams Forget to Hire

8 min read
Team StructureAI EngineeringCareer

Most AI/ML teams I have worked on made the same mistake: they hired for AI and ML, but not for the roles around it, so the AI engineers ended up covering the gap. Ironically, as the ML engineer on most of my projects I did a little bit of frontend, DevOps, and similar work. Back then it seemed like a good decision, but now I realize how inefficient it was for the teams. Even one or two dedicated engineers, backend, DevOps, or frontend, would have sped things up for everyone.

How the Work Lands on AI Engineers

On previous projects most of my colleagues were researchers and data scientists. My part was usually the software engineering: building applications rather than one-off scripts.

Over time that made me the most experienced person with web development and infra, which sounds better than it was. None of those teams had a web developer or a DevOps engineer. It did not occur to me that we were missing a role. My guess is that a team with five AI engineers would have been better off with three AI engineers, one backend engineer, and one DevOps engineer.

When something needed to be deployed, someone had to do it, and the team picked whoever was most likely to get it done. That was usually me, and the decision looked correct to everyone involved, including me. But a dedicated engineer would have done the same work better and faster. My estimate is that what took me a week would have taken an experienced DevOps engineer a day. Nobody was doing anything wrong here, which is why this pattern repeats instead of being one badly run team.

Deployment is only the visible part of this. Most applications also have needs that have nothing to do with AI: monitoring, logging and metrics, user management, roles and permissions, and so on. Someone has to build them, and on an AI team the AI engineers are the closest people to the work, so it lands on them.

At small scale that is fine. At larger scale it is not enough to build a feature that somehow works, it has to work efficiently and keep working as usage grows. That is a specialty of its own, and it belongs to someone who does it every day rather than to whoever happened to be nearest.

And because nothing visibly breaks, the comparison never happens. The feature ships, just later, by a slower route, and there is no moment where anyone notices what it cost.

This is the staffing version of a known problem. Sculley et al. described how little of a real ML system is actually ML code, and how much of it is serving infrastructure, data collection, configuration, and monitoring (Hidden Technical Debt in Machine Learning Systems, NIPS 2015). Their concern was technical debt. Mine is who ends up writing all of it.

It Looks Like Setup, But It Is a Job

At one company I worked for, the machine learning team was new but the company was not. We built internal tools and visualizations to present our results and a few proofs of concept. I ended up doing deployments for the whole team using shell scripts, with no CI pipelines, because I was the closest thing we had to a person who knew how.

What makes this example interesting is that the company already had what we needed. There were full engineering teams developing internal systems and managing DevOps and infra. The support existed. We never asked for it.

I do not think we were being stubborn. The work simply looked like setup, something you do once at the beginning of a project and then you are done with it. If you believe it is a one-off, you do it yourself and move on. If you recognize that it is continuous, you staff it.

Roles that are not building new features, like scripting, maintenance, and keeping things running, are full-time jobs too, because there is always more work coming. A pipeline is not a task that gets completed. It is a surface that needs an owner.

What It Looks Like When the Roles Exist

At Myriad AI, where I work now, I do not write my own deployment pipelines. They already exist. The backend team owns the production surface, the database, the message queue, the monitoring, as their job rather than as a favor to the AI team. When I need something there, I ask the people who own it.

The important word is default. Support is not something I have to know exists and remember to request. It is the normal path, so I never end up quietly absorbing a problem because it did not occur to me that it belonged to someone else.

Where the Line Is

None of this means every team needs specialists for everything.

I have also worked at a company with fewer than five full-time people, where everyone did everything. That was the correct setup and I would do it the same way again. At that size a dedicated DevOps hire is not a missed opportunity, it is a bad idea.

The situation is different once a team is past ten people and has a separate product engineering group. At that size the AI team is not scrappy, it is under-supported, and the two are easy to confuse. One tell is what happens when the team grows. If the answer is always more of the discipline the team already has, the gap does not close, it widens, because every new AI hire produces more work in the boxes nobody owns.

I do not think the threshold is a headcount. It is the point at which the work stops teaching you anything. Building infrastructure for the first time makes you a better engineer. Building it for the fourth time, badly, because there is still nobody else to do it, makes you a slower one.

If I were setting up an AI team today, I would ask three questions. Does the work around the model keep coming back, or was it a one-off? Does someone in the company already own it, and does the team know how to ask? And when the team grows, is it adding only more of what it already has? If the work recurs and nobody owns it, hire for it or give it to a team that already does it, before adding another AI engineer.

Final Thoughts

The reason to hire the roles around AI is not only that the work gets done properly. It is what it does to everyone else on the team.

When someone owns deployment, the AI engineers are not thinking about deployment. When someone owns the frontend, the model work does not pause while a React app gets maintained. The specialist does their own job, and at the same time gives back hours to people who were covering it without ever putting it on a roadmap.