Research / Private equity
A Portfolio-Wide Scorecard for AI Projects
Build a private equity AI scorecard with comparable project outcomes, full costs, quality controls, and finance-verified benefits across portfolio companies.

A private equity operating team needs to know which AI projects deserve more investment, which need repair, and which should stop. Use a shared scorecard that records the same evidence fields across companies, while each management team measures the business outcome that matters to its own operation. Keep realised financial benefits, released capacity, and risk results visible as separate measures.
Start with the decision the scorecard supports
A dashboard with licence counts and a list of pilots tells you how much activity exists. It doesn't tell you whether a project improves a company or whether the next funding request makes sense. Define the decisions first: continue a bounded pilot, expand a working deployment, fix a dependency, or stop a use case that can't meet its acceptance criteria.
Give each project a stable identifier and name the workflow it changes. “AI in finance” is too broad for a project row. “Prepare invoice exceptions for the accounts payable team” lets you connect scope, costs, and results. Record the company, operating owner, finance reviewer, and person who can pause the affected action.
BCG's January 2026 analysis of digital value creation in private equity places digital foundations and AI within portfolio value creation planning. A scorecard can connect that plan to individual delivery decisions. Survey findings about other firms don't establish what your own portfolio will earn from a particular project.
Keep development stages explicit. A discovery project can produce a verified baseline and a testable scope without claiming realised savings. A pilot can establish bounded performance under observed conditions. A deployed project needs operating results and an accountable receiving team. Don't ask a discovery team to invent a return figure just to fill a financial column.
The scorecard below is a suggested operating design, rather than an accounting standard or an industry benchmark. Adapt its fields to the fund's reporting and each company's controls. For the earliest assessment, our guide to turning an investment thesis into an AI opportunity map helps define the work before projects enter the scorecard.
Agree a small set of common reporting fields
Use a common record structure so the operating team can find the same evidence in every submission. Standardise the meaning of the fields, reporting dates, and approval roles. The business metric inside a field can differ: correctly processed orders and resolved support cases represent different jobs.
| Field | Evidence to report | Review question |
|---|---|---|
| Scope and stage. | The workflow, permissions, and live population. | What does this result cover? |
| Business outcome. | The baseline, current result, and sample size. | Did the customer or process improve? |
| Benefit status. | Realised benefit and separately labelled forecasts. | Has finance verified the change? |
| Full cost. | Delivery spend and recurring operating cost. | What does an accepted outcome cost? |
| Quality and risk. | Failures, review effort, and unresolved incidents. | Is expansion within agreed limits? |
| Next decision. | The action, owner, and evidence deadline. | What changes after this review? |
Attach the reporting period, source, and last refresh date to each metric. Show the count behind every rate. A figure of 96% accepted outcomes needs its numerator, denominator, and acceptance rule. Record excluded cases and failed attempts so a team can't improve the apparent score by narrowing the measured population without explanation.
Keep forecast, observed, and finance-verified values in separate fields. A blank field means the evidence is missing; it shouldn't silently become zero. Include a short explanation when the metric doesn't apply at the project's current stage. This makes an early project legible without giving it the same evidence status as a mature deployment.
A single spreadsheet is enough to start if the team maintains definitions and links to controlled evidence. Add a reporting tool when the recurring collection burden justifies it. Avoid making portfolio companies build new integrations solely to populate a board dashboard before the underlying project has a credible baseline.
Keep the business outcome specific to the company
A distributor might measure orders completed correctly by the promised cutoff. A service company might measure tickets resolved without reopening. A manufacturer might measure the time between identifying an exception and an authorised planner deciding what to do. These measures stay close to the work customers and staff experience.
Choose a main outcome and the guardrail that prevents a misleading improvement. Faster support replies need a resolution-quality check. More quotes need a conversion or contribution check. Faster exception handling needs a record of mistakes and downstream rework. Explain which measure drives the investment case and which result prevents expansion.
Use the previous process as the baseline, with comparable volume and case complexity. Record changes in staffing, demand, or service policy during the pilot. If a new workflow receives only easy cases, compare it with the equivalent old population and report the excluded work. Don't present a selected queue as the whole operation.
Where practical, compare equivalent groups or periods under an agreed rollout design. A before-and-after comparison can still inform a decision, but a seasonal surge, new pricing, or a separate process change may explain part of the movement. State those limits. Separate a measured change from a claim that AI caused all of it.
Define acceptance at the end of the job. A draft invoice classification isn't a paid invoice. A generated quote isn't an accepted order. Track review effort, rejected outputs, and work that returns later. Otherwise, the scorecard can reward a project for moving effort into another team.
A corporate group can use the same approach across divisions. A VC platform team can use it to review portfolio pilots, while recognising that it may have less direct access to operating data. In both cases, agree permission to share evidence and use aggregated reporting where customer or employee records don't need to leave the company.
Reconcile benefits with finance
Ask each company's finance lead to approve the bridge from operational result to claimed financial effect. Record the cost line or revenue mechanism, the period, and the evidence of change. Keep the operating team's analysis separate from the company's formal accounting treatment; finance decides how costs and benefits appear in its accounts.
Released staff time can support more throughput or reduce a backlog without reducing payroll. Record that as capacity and show what the team did with it. A reduced overtime bill, cancelled external contract, or approved avoided hire has a different evidence trail. Our article on why time saved isn't the same as money saved explains this conversion in more detail.
For revenue, distinguish additional sales from improved contribution. A larger order book can also create higher fulfilment cost or displace other sales. Identify the baseline, incremental volume, margin assumptions, and what finance can verify. Keep a forecast separate until the relevant outcome occurs.
Cash timing deserves its own line. Earlier collections can release working capital without producing the same amount of recurring operating profit. State whether a claim concerns cash released, expense reduction, or a contribution change. Don't add these figures together under a generic “AI benefit” label.
Here is an illustrative monthly example, not a portfolio benchmark. A document-processing project reduces a verified external processing bill by A$12,000. It adds A$4,000 in recurring software, support, and monitoring costs. The net recurring reduction is A$8,000 per month before any other incremental costs finance identifies. Record the gross reduction and the recurring costs as separate lines so reviewers can reconcile the net figure.
The same project frees 80 staff hours that the company uses to clear a backlog. Keep those hours as a separate capacity outcome. If the external processing bill already reflects work staff brought in-house, check the relationship before adding any labour claim. A second project that uses the same released hours also can't claim an independent saving from them.
For portfolio totals, maintain a benefit register with unique identifiers and an owner for each claimed change. Link dependent projects to the same benefit when they contribute to it. Let the finance reviewers agree the allocation; don't sum two projects' claims just because their dashboard rows differ.
Include delivery and operating costs
Collect delivery costs from the start: integration work, data preparation, security review, staff training, and any parallel operation needed during rollout. Record internal effort as well as supplier invoices, using the company's agreed method. A project can have low software spend while consuming substantial management and engineering time.
Recurring costs include model and infrastructure usage, licences, ongoing support, and the people who review exceptions. Include maintenance when source systems or business rules change. A product-management discussion about post-launch AI costs describes a team receiving bills after development without a clear cost-per-case forecast. It highlights a reporting question, rather than evidence about typical project costs.
Show cost per accepted business outcome with a defined denominator. Count retries and failed attempts in cost even when they don't create accepted outputs. Report volume alongside the unit cost: a monthly platform fee can produce a high unit cost during a small pilot and a lower figure at larger volume. Expansion still needs evidence that quality and support effort hold at that volume.
In the illustrative project above, suppose the company also spent A$24,000 on implementation. At a stable net recurring reduction of A$8,000 per month, simple payback is three months after the full benefit starts. That calculation assumes no further rollout costs or changes in volume. Show the actual delivery period and ramp-up rather than pretending benefits began when the first invoice arrived.
Report spend to date and the cost still required to reach the next decision. A project with substantial sunk cost may still warrant stopping if the next tranche lacks a credible benefit. Conversely, a dependency such as cleaning a shared data source may support several workflows. Record its shared allocation and explain what other projects depend on it.
Keep cost forecasts responsive to workload. Estimate how growth in case volume, longer documents, or more review exceptions changes recurring spend. Use observed usage to update those assumptions. Our guide to the full cost of an AI workflow covers the operating components behind that estimate.
Give quality and risk their own stop conditions
Define quality limits with the company's responsible operating and control owners. A material access failure or unauthorised write needs its own escalation even if the project meets a savings target. Keep the risk status visible alongside the benefit status and prevent an overall weighted score from averaging away a serious incident.
The NIST AI Risk Management Framework's Measure function calls for testing before deployment and regular assessment during operation. It also includes tracking emerging risks. Translate that into a named owner, a monitoring method, and a response when the observed system falls outside the company's approved scope.
Record failures by business consequence and distinguish prevented attempts from completed harmful actions. Report reviewer overrides and cases the system sends to staff. A safe escalation may meet the workflow's design, but it still consumes capacity. The scorecard needs to show both the control outcome and the resulting workload.
Choose thresholds for the particular task before expansion. For a system that prepares internal drafts, review accuracy and correction effort may guide the next release. A system that can update orders needs enforcement of permissions, prevention of duplicate actions, and confirmation of the actual saved state. A generic percentage target can't establish that those controls work.
NIST's March 2026 monitoring report announcement separates questions about functionality, operational service, and impacts after deployment. It also identifies open monitoring challenges. A healthy API and an acceptable model test therefore cover only part of your scorecard; continue checking what happens to users and the business process.
Keep a route for staff and customers to report harm or unexpected behaviour, with a response owner and deadline. Limited incident reports don't prove that incidents are absent if people can't report them. Review detection coverage and unresolved complaints, and record whether an incident requires pausing a specific action or the whole workflow.
Make evidence comparable before aggregating it
Agree reporting dates and the meaning of “current”. A live dashboard refreshed yesterday and an estimate from the first pilot month shouldn't appear as equivalent evidence. Show stale results and missing data explicitly. Preserve the metric definition and its version so a changed acceptance rule doesn't create an unexplained jump.
For rates, report the raw counts and the relevant population. If the same metric applies to several comparable projects, an aggregate rate should use the summed numerator and denominator. An unweighted average gives a ten-case pilot the same influence as a ten-thousand-case operation. Even a correctly weighted total needs project-level results to reveal where failures concentrate.
Don't aggregate unlike outcomes merely because they share a percentage format. Quote conversion and invoice accuracy measure different things. Compare progress against each project's agreed baseline and acceptance criteria, then discuss its benefit and next decision. Keep the local operating measures available to the people who run the work.
For financial reporting across currencies, use the fund's agreed conversion and reporting conventions. State the rate basis and period, and separate recurring annualised estimates from benefits actually realised in the reporting period. Avoid multiplying a good pilot week into a full-year claim without checking demand, seasonality, and the rollout population.
Maintain restricted evidence links behind the summary. Reviewers need a way to reconcile a number with its source, but the portfolio view rarely needs individual customer messages or employee records. Keep each company's data permissions intact and agree how long shared reporting evidence stays available.
Classify evidence in plain language: a forecast using stated assumptions, an observed result with limitations, or a benefit finance has verified. Add the reason when confidence is limited. A label such as “high confidence” without a sample, source, or reviewer adds little to the decision.
Run a review that changes the next action
Ask owners to submit an update that explains changes since the previous period. Start the review with material incidents and decisions that need authority. Then examine outcome gaps, delivery blockers, and requests for expansion. Agree a cadence that matches the risk and project stage; monthly portfolio reporting doesn't replace more frequent operating monitoring.
Use action labels with clear meanings. “Expand” means the current scope met its criteria and a named owner has an approved next scope. “Repair” names the defect or dependency and the evidence needed to retest it. “Pause” states which action stops, who takes over, and what permits resumption. A discovery project can continue to its next evidence milestone without claiming a production success.
Finish each row with a decision, responsible person, and review date. Keep the earlier decisions so the team can see whether a project repeatedly misses the same commitment. If the evidence still doesn't support expansion, revise the scope or stop spending on that use case instead of moving the target without explanation.
Share what another company can actually reuse: a tested control, a supplier assessment, or a reporting definition. Moving a working workflow to a new company still requires its own data mappings, permissions, and acceptance test. The scorecard should make those differences visible while helping the portfolio team direct support to the work with a credible next step.
