For decades, we’ve tried to measure what people know through proxies. Multiple-choice tests, essay prompts, presentations, oral exams. These are all stand-ins for something far more important: what the person can do with what they know.
A teacher can ace an exam by talking about the theory of classroom management, but that doesn’t matter if he loses control of a roomful of fourth graders. Knowledge is the input, but the output that really matters is performance. Until recently, we’ve only been able to measure the former and it has never been enough.
With generative AI, the disconnect is impossible to ignore.
When ChatGPT can craft a competent essay or create a compelling presentation in seconds, the already weak signal from these proxy assessments falls apart entirely. Institutions are responding with AI policies and plagiarism detectors, but these are temporary solutions to a bigger problem. The deeper issue is that the assessments themselves were never measuring the signals that actually matter.
If the most important thing is what a person can actually do, assessments need to observe them doing it. That means watching people work, evaluating the skills and knowledge they use, and then scoring that performance against expert standards. When our lives matter, this is exactly how to do assessment. Think about the hours of clinicals aspiring nurses and physicians complete, under the careful observance of seasoned practitioners. Think about the first officers who spend hundreds of hours in the righthand seat of a cockpit under the watchful eye of an experienced captain. The challenge has always been a question of scale: there aren’t enough experts to evaluate every learner in less high-stakes learning environments.
Multimodal AI could make it possible to change that.
Addressing the challenge of scale
Whether it’s a college student or a new hire, many domains require direct observation to actually deliver a useful assessment. In that aforementioned example, a nurse educator needs to be physically present to evaluate a student’s clinical technique: one observer to one student, one session at a time. The observation that provides the most valuable signal is the one that’s the most expensive and impractical to implement. The result is increased costs for the institution and delayed feedback for the learner, and a huge bottleneck in training more nurses at a time when our healthcare system is desperate for them.
The reason exam and text-based assessment is so prevalent is because it addresses this issue of scale. Dozens or even hundreds of learners can take an assessment at the same time, and an educator (or team of educators) can evaluate the assessment asynchronously. Relevance is sacrificed in favor of logistics.
Multimodal AI bridges the gap. A multimodal AI system is capable of simultaneously interpreting video, audio, and text against a structured rubric. In the same way an instructor evaluates the actions of a student in real-time and measures it against their experience and expertise, multimodal AI can capture and analyze the steps taken by a student in a video recording.
If we return to the example of nursing education, multimodal AI can score a student’s clinical simulation within minutes, assessing everything from hand technique to response sequencing to safety protocol adherence. This is particularly valuable in practice environments. When a student is dependent on an instructor for practical feedback, they could go weeks without receiving useful coaching. With multimodal AI, a student can get scored immediately on a simulation at 10 p.m., identify what went wrong, and then run the simulation again. AI makes assessment faster and more consistent, allowing instructors to focus entirely on the highest-value teaching and assessment.
The impact of time
Multimodal AI could also transform assessment in other areas beyond education and the workforce. When a teacher or parent suspects a child may have ADHD or be on the autism spectrum, the standard of care requires a trained clinician to administer a structured, time-consuming evaluation. The clinician needs to watch the child for an extended period, interpret behavioral signals, and score what they see against a validated rubric.
The supply of qualified clinicians doesn’t meet the overwhelming demand from parents; in many parts of the U.S., families can wait for up to a year to have their child evaluated. During that period, children typically receive no classroom accommodation or formal support. Their learning and development suffers as a result.
A purpose-built AI observational tool could be activated when a teacher or parent has an early concern about a child’s attention or behavior. That tool could surface whether the concerning behavior is persistent, episodic, or absent entirely. For the clinician conducting a formal evaluation, the AI tool becomes an extremely valuable input that allows them to capture multiple structured observations that no single appointment could replicate. The result is that the child receives a diagnosis on a much faster timescale, and they benefit from the necessary intervention significantly earlier.
Proven approaches, new possibilities
It’s worth noting that the concept of performance-based assessment is not new. This type of authentic assessment based on the demonstration of real-world skills has long been used in competency-based education (CBE). The breakthrough surrounds the question of scale. Knowing the value and efficacy of CBE and performance-based assessment, we should pursue every opportunity to increase its use through multimodal AI.
The possible applications are enormous. A language learner gets feedback on tone, pronunciation and pacing, not just grammar and vocabulary. A culinary student is told whether they’re using proper knife techniques. An HVAC apprentice is scored on the precision of a manual procedure on a specific piece of equipment.
There is an entire world of competency that exists beyond written exams. Multimodal AI could make it possible for us to spend more time in that world, practicing and refining the skills that truly matter. From universities to enterprises to culinary schools to the skilled trades, we can finally go beyond assessing what someone knows and measure instead what they can do with what they know.
Arrun is the Managing Director at SJF Ventures and Paul is a visiting professor at the Harvard University Graduate School of Education and the former president of Southern New Hampshire University. They believe that if the most important thing is what a person can actually do, assessments need to observe them doing it (especially in the age of AI).



