The market for AI coaching analysis tools is moving quickly. If you are responsible for coach development at a governing body, a national federation, or a performance programme, you have likely seen more than one of these tools pop up in the past year. Well presented with good websites, polished data dashboards and captivating claims.
The promise of objective data on coaching behaviour, at scale, without being limited by requiring a human observer at every session is real and worth solving. The products being offered to solve it need scrutinised and consideration before jumping in with 2 feet.
Before committing to any AI coaching tool, the following four questions provide a structured basis for evaluation. They follow directly from the two topics already covered in this series:
Questions
1. Can the source data be independently scrutinised?
Request to see the evidence behind a specific classification. If the tool states that a coach asked a certain number of questions in a session, ask to see those moments ie in the transcript, the audio, or ideally the video. A credible provider should be able to demonstrate this in real time during a product evaluation.
If this is not possible, the output cannot be verified and the data, however confidently it is presented, and any subsequent use of it in a coach development decision is built on an unverifiable claim.
2. Is there a mechanism to identify and correct errors?
Ask how the provider handles inaccurate output. Specifically: can a coach or coach developer flag a classification they believe to be wrong? Is that feedback used to improve the underlying model, or does it have no effect on future output?
A tool with no error-reporting mechanism has no structural incentive toward accuracy. It generates reports, distributes them, and no reason to find out whether what it produced was correct. This should be treated as a significant limitation, proceed with caution.
3. What is the tool actually measuring, and what can it not see?
AI can determine that a coach delivered a particular type of behaviour frequently and coherently. It cannot, on its own, determine whether that behaviour was correct, well-timed, or appropriate to the situation. Frequency is not quality and coherence does not mean appropriate, and an organisation needs to avoid accepting these conflations without scrutiny.
Ask directly: is this tool measuring how often a coach did something, or how well? Or how does it know the coach was good at 'x'?
Ask what the tool cannot see at all. A coaching session is not only language. Tactical decisions based on opposition, physical skill execution, non-verbal communication, the score, the coach - athlete relationship and the stage of the season. These are invisible to a tool that processes audio alone. A provider who is clear about these limits is demonstrating a more rigorous understanding of their own product than one who presents its data as a comprehensive coaching insight.
4. Where does the tool stop, and who controls the conversation after that?
A report is not, by itself, a development outcome. Ask what infrastructure exists to support the conversation that should follow: the ability to watch back, annotate, to add context, to link data points to a discussion between coach and coach developer.
Then ask a related but separate question: if the AI-generated report is the judge and jury, who is setting the agenda? The coach responds to what the model noticed. What the coach developer independently observed, or what the coach wants to explore themselves, can get crowded out by the data. The best tools should treat their output as the start of a conversation that the coach and coach developer continue to direct.
If a product's involvement ends at the point the report is generated, the organisation adopting it will need to build that supporting structure itself. That has resourcing implications worth factoring into any decision.
A final consideration
This series opened with the observation that an impressive-looking coaching analysis tool can now be built quickly, and that appearance is not evidence of reliability. Its second part returned to what happened when performance analysis first reached players: accurate data, reduced to a stat, with no room to explain the context behind it. That was a failure of use, not of accuracy.
The same risk now applies to coach development. The organisations adopting AI coaching tools have an opportunity to avoid repeating that history, but only if decisions are based on a structured evaluation of transparency, accountability, and appropriate use.
The four questions above are a starting point for that evaluation. I hope the series provides a useful insight.