The first piece in this series was about accuracy. Can you actually trust what the tool is telling you? Can you verify its output against the source?
This piece assumes you have found a tool that passes that test. The data is as accurate as it can be. The classifications are verifiable. The feedback loop exists.
Accurate data, presented without context and without the right interface and features around it, can still fail coaches. We know this because of what happened to players when video analysis first arrived in sport.
A Problem We Have Lived Before
In the early days of video analysis, data became an endpoint. A player missed three tackles, the data said so and the player was dropped and told to work on their tackling.
The data didn't reveal whether one of those tackles was a chase-back cover effort, while the rest of the team had not tracked back. Whether the misses were a system failure or technical execution. The data was accurate but the analysis provided something too simple to coaches looking for a quick way to make decisions. And it often failed to develop the athlete.
This happened with humans doing the coding. The event data was as reliable as it could be. The problem was never the accuracy, it was the use.
We built Coach Logic 12 years ago as a video analysis platform for teams to specifically push back against that. We made the analysis tools available to players and coaches and built online so everyone had access to the same information.
Instead of data being the endpoint, it surfaced moments, that in turn became the start of a conversation, not the conclusion of one. To give coaches and players something to discuss, not just a verdict to accept.
The risk now is that the same mistakes made in the early days of performance analysis for players is now made when developing coaches. AI-generated coaching reports, presented without context and without the infrastructure for reflection, are the equivalent of "you missed three tackles, you're dropped". Just aimed at coaches instead of players.
We should not make the same mistake twice.
The Context Problem: "It Depends"
Any experienced coach developer knows that the most honest answer to almost any question about coaching is "it depends." :)
It depends on the session intention. The stage of the season. The experience level of the group. The relationship the coach has built with the athletes. What the coach had for lunch...etc etc. The same behaviour, in two different contexts, can be exactly right in one and counterproductive in the other.
A tool that surfaces frequency counts without accounting for context is only providing numbers and should not be mistaken for anything else. The insight requires a human who understands what was happening in that session, on that day, with that group to make sense of it.
The Quality Problem: Frequency is Not Quality
AI can determine that a coach delivered tactical instruction frequently. It cannot determine whether that instruction was correct or appropriate.
Consider a football coach who consistently instructs their team to play with five defenders and exploit space on the counter-attack. The coach communicates this clearly, with good timing. The coach gets a high score for their tactical prowess.
But what if the players are 7 years old? What if the team is losing 3-0? What if there are not five defenders in the squad capable of executing that shape? What if the players do not have the pace to counter effectively? The instruction is being delivered well but it is probably not correct.
In other sports the physical execution of a skill ie a golf swing, a backflip, a dive. The AI can say how frequently technical instruction was provided by the coach but it can't determine its value. Some of the most important dimensions of coaching are invisible to a tool that only processes audio.
Frequency of a behaviour and quality of a behaviour are not the same thing.
The problem arises when tools present frequency data as quality data, and when organisations accept that framing without asking the question.
The Reflection Problem: Who is Setting the Agenda?
When a coach developer sits down with a coach and the AI report is the first thing on the table, it immediately shapes what gets discussed. The conversation follows the data. The coach responds to what the AI thought was significant. The coach developer works through what the model surfaced.
But what if the coach developer observed something the report did not capture? What if the coach had a specific moment they wanted to explore? A decision, a relationship, a response to an unexpected situation that the data did not flag? What if the most important thing that happened in that session was not the thing the model classified with the highest frequency?
A tool that presents its output as the end point is doing something more problematic than being inaccurate. It is substituting the model's judgement for the coach developer's professional expertise and the coach's own reflective instinct. Coach development becomes a process of responding to what the AI noticed, rather than a genuine inquiry into what the coach experienced and what they want to get better at.
The data should support and inform the conversation. Not lead it.
What good infrastructure looks like
The answer to all of this is not to abandon AI in coach development. Phew!
It is to be honest about what AI can and cannot do, and to build the right infrastructure around it. That's been a big part of our mission when developing SAM at Coach Logic.
The data surfaces moments, patterns, raises questions, and provides an objective reference point that neither the coach nor the coach developer could easily generate alone.
But the reflection, the conversation, the professional judgement, the contextual knowledge, is not something the tool provides. That is what the coach developer brings and the tool should support that expertise, not replace it.