Amazing Devs

Managers: Fix AI Team Roles with RACI, NIST RMF & 30/60/90

Managers: Fix AI Team Roles with RACI, NIST RMF & 30/60/90

AI team roles responsibility title card

Every practical AI team needs product leadership, model builders such as research and applied scientists and ML engineers, a platform or MLOps function, and a governance role, and most organizations succeed with a hub-and-spoke model where an AI platform team enforces runtime rules. Smaller teams combine roles; larger ones split them further as usage grows.


TL;DR:

  • A centralized platform team should own runtime infrastructure, model registry, and monitoring tools, while the governance team handles policy and risk classification.
  • Hybrid skill sets are now essential, with candidates expected to bridge research and engineering tasks, including deploying models from notebooks to production.
  • Clear role ownership, explicit SLAs, and mapped tool responsibilities prevent duplication, shadow AI, and accountability gaps in AI governance.
  • Staff augmentation from nearshore providers can quickly fill specific roles like ML engineers or TEVV contractors, accelerating team scaling.
  • Measuring each AI role by specific, relevant KPIs—such as deployment frequency, model quality, or cycle times—improves accountability and team performance.

Amazing Devs
Scale Your AI Engineering Team
Amazing Devs connects businesses with skilled nearshore developers from Brazil, matched for technical ability, cultural fit, and business alignment.

Explore Amazing Devs

Table of Contents

Core AI team roles and responsibilities

Before you post a job listing, know what each role actually owns. Titles vary by company, but the underlying jobs to be done are consistent across the industry.

The AI product manager sets priorities and translates business goals into model requirements. They own the roadmap, define success metrics, and negotiate trade-offs between speed and model quality. Their deliverables include product requirement documents, prioritized backlogs, and stakeholder updates.

A research scientist explores new methods and pushes the boundary of what a model can do. They read and produce research, design experiments, and validate hypotheses before anything reaches production. Expect papers, prototype notebooks, and experiment reports as their output.

The applied scientist sits between research and product, adapting known techniques to a specific business problem, as described in detail for AI-native engineering roles at Opsphere. They design model architectures for a use case, run offline evaluations, and hand off validated approaches to engineering. Their deliverables are trained model candidates and evaluation reports tied to business metrics.

ML engineers turn models into working software. They build training pipelines, optimize inference performance, and integrate models into applications. Their output includes production-ready services, containerized endpoints, and CI/CD pipelines for retraining.

Data scientists focus on the data side of the problem: exploratory analysis, feature engineering, and statistical validation. They often own dashboards, A/B test designs, and feature stores.

Data engineers build and maintain the pipelines that feed everything else. They own data ingestion, transformation jobs, and data quality checks, and their deliverables are reliable, documented pipelines other roles depend on.

The platform or MLOps engineer owns the shared infrastructure: deployment tooling, monitoring, and versioning systems. They build the scaffolding so model builders do not each reinvent serving infrastructure.

An AI platform engineer (sometimes the same person as the MLOps engineer at smaller companies) owns the gateway, model registry, and cost tracking that let multiple teams share infrastructure safely.

TEVV and AI QA specialists test, evaluate, verify, and validate models before and after deployment, a function the NIST AI RMF frames as a continuous, integrated activity rather than a one-time gate. Their deliverables include test suites, bias and drift reports, and sign-off documentation.

The AI governance lead sets policy, classifies risk levels for use cases, and coordinates with legal and compliance. They own the approval process and the audit trail.

UX or context engineers design how humans interact with model outputs, including prompt structures and interface affordances that shape what a model produces. Site reliability engineers (SREs) keep the serving infrastructure available, often shared with the broader engineering organization.

Seniority ladders in these roles usually track two things: scope of ownership (a single model versus a platform serving dozens of teams) and the ability to operate without close supervision. A “senior” applied scientist, for instance, is expected to define their own research direction, not just execute someone else’s.

  • AI/product manager: owns roadmap and success metrics, not model internals.
  • Research and applied scientists: own model quality and validated approaches, not production uptime.
  • ML and platform engineers: own reliability, scaling, and deployment, not research direction.
  • Governance and TEVV leads: own risk sign-off and audit trails, not feature velocity.

The AI roles continuum: when research and engineering blur

Large language models and production-scale training have collapsed much of the old divide between “the people who invent methods” and “the people who ship them.” A roles continuum synthesis documents this shift directly, arguing that research and engineering duties increasingly overlap and that organizations should hire and build ladders around hybrid competencies rather than rigid titles.

In practice, this means a research scientist who cannot containerize a model slows the whole team down, and an ML engineer who cannot read a paper closely enough to reimplement a technique creates the same bottleneck from the other direction. The continuum research recommends treating these as points on a spectrum instead of separate job families, which shortens the loop between an idea and a working system.

For hiring, this changes what you screen for. Instead of asking “is this a research hire or an engineering hire,” ask which side of the workflow is weaker on your existing team and hire toward that gap. If productionizing models is the bottleneck, favor engineering-forward hybrids even for research-titled openings.

  • Look for candidates who can describe both a model’s mathematical assumptions and its deployment constraints.
  • Ask for a recent example where they took an idea from notebook to something another team used.
  • Weight systems thinking as heavily as modeling depth for mid-level and senior hires.

Pro Tip: Write job postings around outcomes (“productionize retrieval models”) instead of titles (“Research Engineer II”) to attract candidates who fit the continuum rather than a narrow label.

AI platform team vs AI Center of Excellence: who owns what

Confusing these two functions is one of the most common ways AI initiatives stall. An AI platform team playbook draws the line clearly: the platform team owns runtime infrastructure, while a Center of Excellence (CoE) owns policy and use-case approval. Conflating the two is a recurring scaling mistake.

The platform team’s job is operational: it runs the AI gateway that routes and rate-limits requests, maintains the model and agent registry that tracks what is deployed and its risk classification, builds observability dashboards, and manages cost through FinOps practices. The CoE’s job is normative: it sets the rules for what counts as an acceptable use case, classifies risk tiers, and approves or rejects proposals before they reach the platform.

A practical handoff sequence looks like this:

  1. The CoE reviews a proposed use case and assigns it a risk classification.
  2. The CoE documents required controls (human review, monitoring thresholds, data restrictions) based on that classification.
  3. The platform team provisions access through the gateway and registers the model with its classification attached.
  4. The platform team enforces the controls at runtime, gating access until requirements are met.
  5. Both teams review flagged incidents together, with the platform team supplying monitoring data and the CoE deciding on policy consequences.

The risk of conflation is concrete: a CoE that writes standards with no enforcement hooks into the platform produces shadow AI, because teams route around a policy that nobody can technically check. The fix is to keep the CoE scoped to policy and require platform gating, meaning gateway and registry checks, before any team gets runtime access. Platform teams only enable continuous governance when their tooling actually enforces the policy; otherwise the CoE’s rules stay aspirational.

Common org patterns, sizing, and staffing benchmarks

Three organizational patterns cover most companies building AI capability. A centralized model puts all AI talent in one team serving requests from the rest of the business, which works well early on but creates a bottleneck as demand grows. A federated model embeds AI talent directly in product teams, which speeds delivery but risks duplicated infrastructure and inconsistent governance. A hub-and-spoke model, recommended by the platform playbook, keeps a central platform team owning shared infrastructure and governance enforcement while spoke teams embed model builders close to product work.

Staffing benchmarks suggest approximately one platform engineer per several model builders after a team has scaled beyond initial use cases. Early-stage platform teams often start with a small group and scale up significantly as they begin serving multiple business units.

  • Add a dedicated platform engineer once more than two product teams depend on shared model infrastructure.
  • Add a TEVV specialist once models touch decisions with legal, financial, or safety consequences.
  • Add FinOps capacity once model inference costs become a line item leadership tracks monthly.
  • Evolve from centralized to hub-and-spoke gradually: keep governance central while spinning off embedded builders team by team, rather than reorganizing everyone at once.

Governance, RACI, and accountability for AI risk management

Governance work fails when everyone assumes someone else owns it. A RACI-style breakdown, mapped to the activities the NIST AI RMF frames as core to trustworthy AI, gives every role a clear answer.

  • Risk mapping: the AI governance lead is accountable, the product manager and applied scientist are consulted, and the platform team is informed.
  • TEVV: the TEVV specialist is accountable, the applied scientist and ML engineer are responsible for supplying test artifacts, and the governance lead is consulted on thresholds.
  • Model registry maintenance: the platform engineer is accountable and responsible, with the governance lead consulted on risk classification fields.
  • Continuous monitoring: the platform team is responsible for the tooling, the TEVV specialist is accountable for interpreting drift signals, and the SRE is informed for availability issues.
  • Incident response: the platform team is responsible for containment, the governance lead is accountable for disclosure decisions, and the product manager is informed.
  • Access controls: the platform engineer is responsible for implementation, the governance lead is accountable for policy, and all model builders are informed.

TEVV and continuous monitoring only work when they are wired into platform tooling rather than handled as a separate spreadsheet exercise. The NIST framework treats TEVV as an integrated, ongoing activity, which means monitoring dashboards, drift alerts, and incident playbooks need to route directly into the same systems the platform team already runs.

Pro Tip: Document every governance handoff as an explicit SLA, for example “risk classification within three business days,” so accountability does not quietly drift back onto whichever team is easiest to blame.

Hiring, onboarding, and a practical checklist for building an AI team

A structured hiring process catches hybrid gaps before they become expensive. Test for code, modeling judgment, systems thinking, and TEVV awareness together rather than in separate, disconnected rounds.

  1. Run a notebook-to-production exercise: give candidates a prototype model and ask them to containerize and serve it, which reveals both ML understanding and engineering discipline.
  2. Run an incident simulation where candidates interpret monitoring signals and propose a remediation plan, testing operational judgment under realistic conditions.
  3. Include a short TEVV checklist exercise where candidates identify what should be tested before a model ships.
  4. Score against a rubric that weighs systems thinking as heavily as modeling depth, consistent with what AI lead job postings commonly expect from senior technical hires.

Onboarding works best on a 30/60/90 structure: the first 30 days focus on shipping a small, low-risk change to build context; days 31 to 60 add ownership of a real backlog item; by day 90, the hire should be running an independent experiment or feature with light supervision. Soft skills matter as much as technical depth here, since collaboration and cultural fit determine how quickly a new hire integrates into existing rituals like standups and code review.

When bringing in nearshore ML engineers or platform specialists, give them clear product-side ownership from day one and fold them into the same sprint cadence as in-house staff rather than treating them as a separate workstream.

How Amazing Devs helps teams scale AI-capable engineering quickly

Nearshore developers can fill roles such as ML engineers, data engineers, platform and MLOps engineers, and QA contractors who support TEVV work. The process should center on assessing both technical skill and cultural fit before a placement, then managing the contract and administrative work so the client team can focus on integration.

Before starting a pilot engagement, confirm the screening steps used for a given role, agree on a 30/60/90 ramp plan matching the onboarding approach above, and clarify who owns outcomes on the product side. Amazing Devs handles staff augmentation and team extension engagements built around these same integration principles.

Tools and technologies tied to each AI role

Tool ownership should map directly to the responsibilities above, so nobody assumes a tool is “someone else’s problem.” Research and applied scientists typically work in notebook environments and experiment trackers to manage model iterations. ML engineers own the training and serving stack: containerization tools, model-serving frameworks, and CI/CD pipelines for retraining. Data engineers own pipeline orchestration and warehouse or lakehouse platforms.

Platform engineers own the AI gateway, model and agent registry, and observability dashboards described in the platform playbook, along with cost management tooling for tracking inference spend. TEVV specialists own testing frameworks for bias, drift, and robustness checks, often layered on top of the same observability stack the platform team maintains. Governance leads typically own documentation and audit-trail systems rather than technical tooling directly, since their job is tracking decisions and approvals rather than running infrastructure.

AI roles mapped to tool ownership

Assigning a named owner to each tool category, rather than leaving it as a shared responsibility, prevents the common failure where monitoring dashboards exist but nobody checks them. A tool without a clear owner tends to get built once and abandoned.

KPIs that show whether AI roles are working

Each role needs a metric tied to what it actually controls, not a generic team-wide number everyone reports the same way. Product managers can be measured on roadmap delivery against committed timelines and on adoption of shipped features. Research and applied scientists are better measured on model quality metrics relevant to the use case, plus how often their work reaches production rather than staying in a notebook.

ML and platform engineers are usually measured on deployment frequency, inference latency, and system uptime, since those numbers reflect the infrastructure they are accountable for. TEVV specialists can be measured on test coverage of deployed models and the time between a detected drift signal and a resolved incident. Governance leads are best measured on cycle time for risk classification and approval, since a slow approval process is often the first sign that policy and platform are not talking to each other.

Avoid a single shared KPI dashboard that blends all of these together. When every role reports the same top-line number, it becomes impossible to tell which part of the team is actually underperforming, and accountability quietly disappears.

How AI teams work with product, engineering, and compliance

AI teams rarely succeed in isolation, and the handoffs to other business units matter as much as internal role clarity. With product teams, the connection usually runs through the AI product manager, who translates business priorities into model requirements and reports model performance back in terms product leadership can act on. This is also where AI-native development practices reshape how product and engineering plan sprints together, since model behavior can shift in ways a traditional feature never does.

With broader engineering, the platform team’s shared infrastructure, including the gateway and registry from the earlier section, becomes the connective tissue, since other engineering teams depend on the same deployment and monitoring tools rather than building their own. Compliance and legal typically connect through the governance lead, who translates regulatory obligations into risk classifications and approval criteria the rest of the team can act on without needing legal training themselves.

Regular touchpoints work better than ad hoc escalation: a standing review where governance, platform, and product each report status prevents the common failure where compliance only hears about a model after it ships. Applied scientists and ML engineers benefit from sitting in on at least some of these reviews directly, since translating governance requirements through two layers of management tends to lose the technical detail that actually matters for implementation.

Common pitfalls managers repeat when standing up AI teams

The most damaging mistake is conflating platform and CoE responsibilities, which leaves policy unenforced and infrastructure duplicated. A close second is hiring for rigid titles instead of the continuum of skills a team actually needs, which stalls production timelines. Treating TEVV as an afterthought rather than a continuous discipline is the third. Fix role clarity and enforcement hooks before headcount. If your team needs to scale fast without sacrificing that clarity, a staffing partner with a rigorous vetting process is worth considering.

— Gabriel

How Amazing Devs can help you hire AI-capable engineers fast

Once your roles and governance model are clear, the bottleneck is usually finding people who fit them without a months-long search. Nearshore developers from Brazil can be sourced for these roles, screened for technical skill and cultural fit before interviews, and contract and administrative work can be handled to speed international hiring.

Amazing Devs

  • Nearshore staff augmentation can fill specific gaps, like an ML engineer or platform specialist, alongside your existing team.
  • Team extension models can embed a group of nearshore engineers directly into your product workflow for ongoing capacity.
  • Nearshore QA contractors can support the TEVV and testing work that governance sections above depend on.

A pilot engagement typically starts with a scoped role, a defined ramp plan, and clear ownership on your side, the same structure recommended earlier in this guide. Check current openings through nearshore staff augmentation or explore the Team Extension Model if you need ongoing embedded capacity.

Sources

FAQ

What is the 30% rule for AI?

There is no established “30% rule” defined in AI governance or staffing frameworks like the NIST AI RMF or major platform playbooks. If you encountered this term in a specific context, it likely refers to an internal benchmark from that source rather than an industry-wide standard.

What types of AI roles are there?

AI roles generally span product leadership, model building (research scientists, applied scientists, ML engineers, data scientists), platform and MLOps engineering, and governance functions like TEVV specialists and AI governance leads. Many organizations now treat these as points on a continuum rather than fixed titles, since research on hybrid AI roles shows responsibilities increasingly overlap.

What are the 7 types of AI agents?

This guide focuses on human team roles rather than technical agent architectures, so a specific taxonomy for AI agents falls outside what it covers. Definitions of agent types vary by source and technical context, so check a dedicated AI architecture reference for that breakdown.

What is the typical structure of an AI team?

Most successful AI teams use a hub-and-spoke structure, where a central platform team owns shared infrastructure and governance enforcement while smaller groups of model builders embed with individual product teams. This pattern, described in platform team playbook guidance, balances consistency with the speed that embedded teams need.

How do I start building an AI team if I have no existing staff?

Start with a small hub covering product, one or two model builders, and a platform or MLOps generalist, then add governance and TEVV roles as use cases touch higher-risk decisions. Nearshore staffing options like Amazing Devs’ staff augmentation can fill specific role gaps quickly while you build out a permanent structure.