Enterprises are generating code faster than ever, yet few are shipping faster. GitLab’s data show that 78% of organizations see faster commits, but only 21% report productivity gains across the full lifecycle. The gap, argues Manav Khurana, Chief Product and Marketing Officer at GitLab, lies in trust. Every AI-written change still needs review, testing, security checks, and approval, and manual handoffs turn saved time into queues.
In this conversation with CIO&Leader, Khurana explains why prompt logging is overrated, how agents should work under a delegate’s permissions, and why “cost per accepted change” beats commit volume as a metric. He also unpacks the “shadow software factory” risk, India’s evolving software supply chain expectations, and why security is the next layer that needs AI-native governance.

CIO&Leader: GitLab’s data shows 78% of organizations see faster code commits, but only 21% see full-lifecycle productivity gains. Where in the pipeline is that gain getting lost?
Manav Khurana: As code generation increases, the biggest constraint becomes trusting it. Every code change still needs review, testing, security checks, remediations, and approval. That delivery process often relies on manual handoffs and workflows that struggle to keep pace with the additional volume. That’s why the time saved writing code turns into a queue. Review queues grow faster than reviewers can clear them, test suites become a maintenance burden, and scans find more vulnerabilities than teams can fix.
Most organizations find code generation is already the smallest part of the full software delivery process. Organizations need to redesign the broader software delivery system if they want faster code commits to drive faster business outcomes.
CIO&Leader: When AI agents generate a large share of a codebase, how should code review workflows be re-architected, since traditional peer review wasn’t built for that volume or velocity?
Manav Khurana: Code reviews are the obvious new bottleneck with the volume of code that agents produce. Every team has to shift from reading every line to deciding which changes need a person.
Both our engineering teams and our customers now automatically invoke Duo Agents for code reviews on every merge request. The Duo code review flow takes minutes to complete a review of changes against the team’s policies, flag risks, and prompt fixes. Without it, this task can take hours to do and sometimes days of waiting in handoffs.
Our first-hand experience with agentic code reviews has led to an updated policy, in which low-risk changes can proceed with agentic code reviews after they pass the required checks. In contrast, higher-risk changes are automatically escalated to a human reviewer. Either way, both agentic and human code reviews leave a record of all their work in the same system of record for future auditing and context.
CIO&Leader: What does “traceability” mean at a technical level for AI-generated code? Does it mean commit-level attribution, model versioning, prompt logging, or all three?
Manav Khurana: Traceability includes all three and more, but they aren’t equally useful.
The record that matters most is what the agent did: the files it read, the commands it ran, the credentials it used, the systems it accessed, and the approvals it received, in the context of the task or workflow for which the agent was invoked. That allows platform teams to audit agent work on demand and surface risk or anomalies.
Model versioning matters because enterprises may run several models at once and replace them as better ones arrive that provide better price-performance ratios for different tasks. This is a great way to manage costs as agentic use scales up.
Prompt logging is the one thing teams overinvest in. At agent scale, nobody reads stored prompts. In an incident, you need the action record. Capture it as the work happens, on the same path for people and agents – connecting the original work item to the merge request, testing, security checks, approval, deployment, and production outcome.
CIO&Leader: If an AI agent introduces a defect in production, how does GitLab’s platform help trace that defect back to the specific agent, prompt, or pipeline stage that produced it, and verify trust in the result?
Manav Khurana: In the GitLab platform, agent actions are logged against an identity and policy as they run, connecting the change to the agent, the models and tools involved, the permissions it held, the policies that applied, the approvals it required, and the checks that ran.
GitLab Orbit’s context graph then connects code, work items, pipelines, security findings, deployments, and production signals. That gives engineers a path from a production issue back through the development lifecycle, so they see not only what changed, but why it changed and what happened afterward.
To verify the result, teams can check the agent’s tool calls against the policy that governed the change. Once the defect is understood, the fix follows the same path and is tested against the relevant code, configuration, and deployment path before it reaches production.
CIO&Leader: How does GitLab’s architecture bring DevSecOps into a single platform to close the gap between fast code generation and fast, safe shipping?
Manav Khurana: GitLab closes the gap by bringing together the tools for every stage of the software development lifecycle from planning to shipping into one connected platform. Most organizations started by applying AI to code generation. Now, testing, review, security, and deployment have to handle the additional volume.
GitLab connects code, project plans, pipelines, security scans, compliance checks, and deployment history in a single data model, so people and agents share the same context. An agent reviewing a change sees the same pipeline results, security findings, and history a person would.
GitLab organizes that infrastructure into four layers on one platform:
1. Orchestration layer – to coordinate agents that pick up work across every stage of the software lifecycle.
2. Context layer – a graph of code, merge requests, work items, pipelines, deployments, vulnerabilities, ownership, and production signals.
3. DevOps layer – for source control, CI/CD, and artifact management workflows that can handle the higher volume of branches, retries, and pipeline runs that agents create.
4. Security layer – to detect, triage, and fix vulnerabilities in new code at machine scale, as well as apply identity, policy, approval, and audit to both agent and people actions.
CIO&Leader: What technical guardrails should engineering teams put around AI coding agents before letting them commit directly to production branches?
Manav Khurana: Treat each agent as a delegate that can never do more than the person it works for.
In GitLab, an agent runs under a composite identity. Its actions are attributed to the user, and it is limited to the permissions of the user who invoked it. That tells you which actions an agent took, and no agent can escalate privilege. Secrets follow the same rule: Secrets Manager releases each credential only to the job that needs it, based on the environment, the branch, and whether that branch is protected. The policy applies wherever the agent runs, and high-risk actions stop for approval that records who approved.
Test the boundary before granting more autonomy. Can the dependency agent retrieve a production credential? Can an agent remove a failing security job and merge? If yes, it isn’t ready for production branches.
CIO&Leader: How does fragmented tooling separate CI/CD, security scanning, and AI code-gen tools, and how does that technically compound the governance problem you’re describing?
Manav Khurana: Every tool holds a piece of the record, and none holds the whole change.
Plan in Jira, review in GitHub, build in Jenkins, scan with Snyk, store in JFrog, deploy with Harness. That is six identity models and six logs. At agent volume, stitching them together becomes a bottleneck in itself, and no single policy covers every stage and every tool.
That is a shadow software factory. It runs on personal laptops, with each developer orchestrating agents according to their rules, with local credentials and unmanaged tools. Standards are their choice. There is no evidence chain from plan to production. This is what leads to unintended risks to reliability, security, and privacy, as well as ballooning AI costs.
The antidote is a governed software factory. It’s where the standard for how agents handle the software development lifecycle is defined on the DevOps platform. Agents are orchestrated consistently. Policies are enforced centrally, producing a connected evidence chain from plan to production. And we end up measuring the team outcome of shipping accepted changes and the cost per accepted change.
CIO&Leader: Beyond commit volume, what technical metrics should engineering leaders track to measure real AI impact: cycle time, defect escape rate, and review latency?
Manav Khurana: I’d start with cost per accepted change. It’s the total cost of people, AI, and tooling for planning, building, reviewing, securing, and deploying software, divided by the number of changes that reach production.
Commit volume and AI usage show activity; cost per accepted change measures what shipped. Retries and abandoned work add cost without adding to the count, so waste shows up in the number.
The other metrics help show why an accepted change costs more or less. Cycle time shows how long it takes to move from idea to production. Review latency shows where work is waiting. Defect escape rate, change failure rate, and mean time to recovery show whether faster output is creating more risk or rework. Deployment frequency shows whether more code is actually reaching users.
Teams should also track how quickly security issues move from detection to verified remediation, and whether changes adhere to the required controls. Together, these measures show where accepted changes are becoming slower, more expensive, or riskier.
CIO & Leader: With Indian regulators pushing for stronger software supply chain accountability, what technical capabilities does an enterprise need to demonstrate provenance and trust in AI-generated code?
Manav Khurana: Indian regulators already ask firms to account for what is in their software. SEBI’s cyber resilience framework requires regulated firms to keep a software bill of materials for critical systems. CERT-In’s guidelines recommend one that records the author of each component, and its 2025 update adds a bill of materials for AI.
To prove it, an enterprise needs an identity for every agent and an inventory of the models and dependencies it uses. It also needs signed commits and artifacts, as well as policy and security checks that run automatically in CI/CD. Every change requires a tamper-evident audit record, with a named person approving high-risk changes. With these in place, you can prove the software in production is exactly what was reviewed. The upcoming GitLab Security Standard calls this the Release control.
CIO&Leader: Looking at GitLab’s roadmap, what’s the next technical layer after code generation that needs AI-native governance: testing, security scanning, or deployment approval?
Manav Khurana: All three, but security is most prescient. Models find vulnerabilities faster than teams can fix them. According to research from Palo Alto Networks’ Unit 42, a single AI system uncovered more than 14,000 vulnerabilities across almost 4,000 open-source projects in two months, with 99% of them never having been reported before.
Our platform brings scanning, remediation, testing, and deployment approval into a single governed loop. Agents fix routine findings, and people approve the high-risk ones.
We are documenting the controls, architecture, metrics, and practices teams need to establish a baseline, and we will go deeper on all of it at GitLab Transcend on October 6.
