PreviewYou're previewing CTO AI Forge. Sign in to use it.Sign in
← Tools

Signature tool · CTO AI Forge

AI engineering productivity scorecard

Log one row per team per period — AI adoption, PRs merged, engineers, cycle time, change-failure rate, review time, AI-authored share and deploys — by hand or by pasting a CSV. The scorecard computes period-over-period trends, PRs per engineer and DORA-style bands, and a rule-based diagnosis of the limiting factor (for example "review is the bottleneck" or "tighten quality gates"). Published benchmarks are shown with their sources; the rules and thresholds are illustrative and editable by your team's judgment. The math runs in your browser; AI only explains the result.

1 · Log your teams

Example data — four fictional teams over two quarters, so you can see how the scorecard works. Replace them with your own numbers, or paste a CSV below.

Team periods: one row per team per period
TeamPeriodAI adoption %PRs mergedEngineersCycle time (h)Change-failure %Review time (h)AI-authored %DeploysRemove

Cycle time: commit → production if you have it (that is what DORA’s lead-time bands describe). Review time: PR opened → approved. Percentages are 0–100.

Paste CSV from your tools

First row = header. Recognized columns: team, period, AI adoption %, PRs merged, engineers, cycle time (h), change-failure %, review time (h), AI-authored %, deploys (common variants like “cfr” or “deployments” work too). Comma, semicolon or tab separated.

2 · Scorecard

Teams
4
Engineers
29
PRs / engineer
25.6
AI adoption
68.3%
AI-authored
27.1%
Change-failure
8.1%

Latest period of each team. Adoption and AI share are engineer-weighted; change-failure rate is deploy-weighted.

Payments

2026-Q1 → 2026-Q2
Limiting factor
Review is the bottleneck

AI adoption is 82% but PRs merged per engineer changed only 3.1%, and review time is 27 h (+92.9%).

Next step: Code is being written faster than it can be reviewed. Cap PR size, use AI-assisted first-pass review, set review SLAs and spread review load beyond a few senior engineers.

AI adoption
82%+27 pts
PRs / engineer
26.8+3.1%
Cycle time
58 h+26.1%
High (1 day – 1 week)
Review time
27 h+92.9%
Change-failure
10%+1 pts
High / medium range (10–20%)
Deploys / week
8-5.5%
Elite (on demand, daily or more)
AI-authored
31%+13 pts

Mobile

2026-Q1 → 2026-Q2
Limiting factor
Quality gates

Change-failure rate rose 7 points while the AI-authored share rose 18 points (2026-Q1 → 2026-Q2).

Next step: Tighten quality gates before adding more AI output: required tests on AI-written changes, security scanning in CI, smaller PRs and a clear rule for which changes need senior review.

AI adoption
78%+18 pts
PRs / engineer
26.9+19.4%
Cycle time
66 h-5.7%
High (1 day – 1 week)
Review time
13 h+8.3%
Change-failure
15%+7 pts
High / medium range (10–20%)
Deploys / week
3.2+5%
High (daily – weekly)
AI-authored
38%+18 pts

Platform

2026-Q1 → 2026-Q2
Good sign
Gains are showing

PRs merged per engineer rose 24% without a matching rise in change failures.

Next step: Keep measuring; check the gain holds over two more periods and that PRs are not just getting smaller.

AI adoption
58%+18 pts
PRs / engineer
31+24%
Cycle time
26 h-13.3%
High (1 day – 1 week)
Review time
9 h0%
Change-failure
5%-1 pts
Elite level (≈ 5%)
Deploys / week
13.1+13.3%
Elite (on demand, daily or more)
AI-authored
22%+10 pts

Data

2026-Q1 → 2026-Q2
Limiting factor
Delivery foundations (CI/CD)

Developers use AI, but cycle time is 240 h and only 0.8 deploys per week. Faster coding cannot get through a slow pipeline.

Next step: Fix the foundations first: faster CI, trunk-based development or small batches, automated tests and deploys, before buying more AI tooling.

AI adoption
38%+6 pts
PRs / engineer
14.4+2.9%
Cycle time
240 h-7.7%
Medium (1 week – 1 month)
Review time
18 h-10%
Change-failure
11%-1 pts
High / medium range (10–20%)
Deploys / week
0.8+11.1%
Medium (weekly – monthly)
AI-authored
8%+2 pts
Trend: PRs merged per engineer
08.517.125.634.12026-Q12026-Q2
Chart data as a table
Team2026-Q12026-Q2
Payments2626.8
Mobile22.526.9
Platform2531
Data1414.4

How the bands and diagnosis work

The band edges below follow the ranges DORA publishes for its performance clusters; the sourced DORA table appears here once it is published.

  • Cycle time bands: Elite (< 1 day) ≤ 24 h · High (1 day – 1 week) ≤ 168 h · Medium (1 week – 1 month) ≤ 720 h · Low (1 – 6 months) ≤ 4,380 h.
  • Deploy frequency bands: Elite (on demand, daily or more) · High (daily – weekly) · Medium (weekly – monthly) · Low (monthly – every 6 months).
  • Change-failure bands: Elite level (≈ 5%) ≤ 5% · High / medium range (10–20%) ≤ 20% · Low level (≈ 40%) ≤ 40%.
  • Band edges follow the ranges in DORA's 2024 cluster table (see the DORA items). DORA's clusters are survey results, not standards; the band edges here are our simplification for comparison.

Diagnosis rules illustrative assumption — use your own numbers

  • Quality gates: change-failure rate up ≥ 2 pts while AI-authored share up ≥ 5 pts.
  • Review is the bottleneck: adoption ≥ 50%, PRs per engineer up ≤ 5%, and review time ≥ 24 h or up ≥ 20%.
  • Delivery foundations: adoption ≥ 30% but cycle time ≥ 168 h or fewer than 1 deploy per week.
  • Adoption: adoption below 30%.
  • Gains showing: PRs per engineer up ≥ 10% without a change-failure rise.

The diagnosis thresholds are illustrative assumptions, not published benchmarks — adjust them to your organization. With only one period, the scorecard can flag “watch” items but not trends. Use it at team level only — never to rate individuals.

Sign in to save your work privately and come back to it. Export works without an account.

The scores are calculated in code. AI only explains them — check anything important.
The other signature tool →⚖️ Build-vs-buy & model selection matrix