One Witness Isn’t Enough: How to Build Evaluation Data People Actually Trust

L&D has a credibility problem, and it starts with what the field is willing to accept as proof. For years, most organizations have run training evaluation like a defendant grading its own exam: pick the friendliest available number, usually a satisfaction score, and call it evidence the program worked. Boards and CFOs have gotten less patient with that arrangement, not more. The function that wants a seat in the business conversation has to stop presenting testimony and start presenting a case — and that starts with two disciplines most evaluation practice still treats as optional: validity and triangulation.

We’ve Been Grading Our Own Homework

Let’s look at an example. Consider a retail chain that runs a year-long customer service program. Ninety-eight percent of employees rate it excellent. Completion is high. Everyone feels good about it. Then the quarterly numbers arrive: complaints haven’t dropped, returns are still high, and customer satisfaction hasn’t moved. The survey wasn’t dishonest. It just never measured the thing that mattered, because it was the only thing anyone bothered to measure. This is the default posture of training evaluation across most organizations, and it survives for an uncomfortable reason: a single flattering metric is easy to produce, easy to present, and easy for no one to challenge — until the business results fail to follow.

Validity Is the Line Between Data and Noise

Validity is a simple, unforgiving test: does the evaluation actually measure what it claims to measure, or just what was easiest to collect? Asking someone whether they enjoyed a cooking class tells you nothing about whether they can cook a great meal. Liking something and doing it well are not the same claim, and the field has spent decades quietly treating them as interchangeable — tracking satisfaction, completion, and attendance, then asking those numbers to forecast performance they were never built to predict. Evaluations lose validity in a small number of predictable ways: an assessment that doesn’t reflect the real objective, data collection that’s biased or inconsistent, stakeholders with too little at stake to answer honestly, and a report that lets one metric carry the entire argument. Any one of these is enough to produce a confident, well-designed report that is simply measuring the wrong thing, and confident wrong answers are more dangerous to a function’s credibility than honest uncertainty.

Borrow the Method: How a Case Gets Built

The fix isn’t more surveys. It’s triangulation — a discipline evaluation should have borrowed from investigative journalism and law long ago. No credible reporter runs a story on one source, and no credible evaluation should rest on one either. Triangulation strengthens a finding by pulling from multiple data sources, methods, and perspectives until independent lines of evidence agree. There are four distinct types worth naming precisely, because each closes a different gap. Data triangulation cross-references self-reports, peer feedback, and manager observation to see if they confirm the same result. Methodological triangulation checks a survey, a focus group, and a hands-on skills check against each other to see if they land on the same outcome. Stakeholder triangulation gathers new hires, HR, and supervisors so no single vantage point gets mistaken for the whole picture. Temporal triangulation checks whether a result holds up weeks or months later, rather than only in the glow immediately after a program ends. One data point is an opinion. Four independent ones that agree are a finding — and a finding is what a board will act on.

The Model Already Solved This

This isn’t a new layer bolted onto the Kirkpatrick Model — it’s the discipline the four levels were built to enforce. A credible evaluation strings evidence across all four: Level 1, reaction and relevance; Level 2, learning, confidence, and commitment, validated through real skills demonstrations rather than a quiz; Level 3, behavior on the job, validated through manager observation, dashboards, and follow-up interviews; Level 4, business results, validated through ROI, trend analysis, and benchmarks. Pulling data from multiple sources, using multiple methods, across all four levels is what Kirkpatrick calls blended evaluation. Executives were never actually asking L&D for more numbers. They were asking for the ability to trust the ones already on the table — and trust is a triangulation problem, not a dashboard problem.

The Proof Is in the Data That Disagreed With Itself

A financial services firm training frontline managers on leadership skills initially tracked only engagement surveys and knowledge checks — Level 1 and Level 2 data. Satisfaction was high. Performance and retention weren’t moving, and no one could explain the gap using the data on hand. Once the firm added Level 3 and Level 4 evidence — manager observation logs, peer feedback, 360 reviews, business metrics — the story changed entirely: increased collaboration and problem-solving, a 15% lift in engagement scores where extra coaching was added, a 12% reduction in voluntary turnover, and a 9% increase in client satisfaction. No single measure carried that case. Four independent ones, pointed in the same direction, did — which is precisely the standard the rest of the business already holds itself to, and the standard L&D has to meet to be trusted at the same table.

What the Next Decade of L&D Will Demand

The organizations getting real influence out of their training investment aren’t the ones with the most sophisticated dashboards. They’re the ones whose data survives being questioned. That bar is only going to rise as L&D budgets face more scrutiny, not less, and as more of the profession competes for the same credibility that finance and operations have long taken for granted. The practical move is a short audit, run on every program currently in flight: where is this data actually coming from, one source or several? Which of the four types of triangulation — data, method, stakeholder, or time — is already covered, and which is conspicuously missing? Is the number carrying the most weight in the current report also the one most exposed to bias? The goal was never to collect more data. It’s to collect enough independent evidence that the story holds up under scrutiny — because credible evaluation doesn’t just measure impact. It builds influence, and influence is what the function actually needs.

Validity and triangulation solve how to trust the evidence. There’s a bigger question still ahead: how evaluation becomes part of how an organization operates day to day, rather than something done only at the end of a program. That’s the subject of the next episode in this series on the Kirkpatrick Model.

It’s still a great time to pick up Building a Culture of Evaluation and follow along before the next module.

Join the book club here: https://lp.constantcontactpages.com/sl/lSZFbHe

Grab a copy of the book here: https://www.amazon.com/Building-Culture-Evaluation-Kirkpatrick-Performance/dp/1963392337/

Watch the podcast now: https://youtu.be/61PGV_Lyd3Y