Would You Let an AI Coach Your Employees?

AI roleplay tools could make employee training more accessible and repeatable, but automated scoring, workplace monitoring and employment use require careful governance.
Corporate AI video platforms are beginning to move beyond passive training content. Instead of watching an avatar explain how to handle a difficult conversation, employees can now speak to an interactive avatar, practise the conversation and receive automated feedback.
Synthesia, known for AI-generated business videos, has introduced a product called Roleplay Sessions. Synthesia describes it as allowing employees to rehearse workplace scenarios — including sales conversations, customer complaints, management discussions and performance conversations. The system can reportedly respond dynamically and assess the interaction against an employer-defined rubric.
That could make practice more available and less intimidating. But it also introduces important questions:
- Who decides what “good” communication looks like?
- Can the scoring be trusted?
- Are transcripts retained?
- Can managers see individual results?
- Could the data later influence appraisals or disciplinary decisions?
- Could certain accents, disabilities or communication styles be disadvantaged?
- Is the tool truly coaching — or is it employee monitoring?
The central risk
AI coaching could give employees a safe place to practise, but only if the organisation prevents the practice room from quietly becoming an assessment centre.
Last checked: 29 July 2026.
What has Synthesia launched?
Synthesia is known for AI-generated video content used in corporate training. Roleplay Sessions represents a move from passive video — where an employee watches an avatar explain a process — into interactive practice, where the employee participates in a simulated conversation.
| Status | Detail |
|---|---|
| Confirmed | Synthesia has announced Roleplay Sessions as a product for practising workplace conversations |
| Confirmed | The system involves AI avatars that can respond during a simulated conversation |
| Synthesia describes | Scenarios including sales, customer service and management conversations |
| Synthesia describes | Automated feedback and scoring against an employer-defined rubric |
| Synthesia describes | Analytics visible to the organisation |
| Not publicly confirmed | Full extent of transcript retention and storage |
| Not publicly confirmed | Whether facial or emotion analysis is used |
| Not publicly confirmed | Specific default manager visibility settings |
| Not publicly confirmed | Complete subprocessor list |
| Not publicly confirmed | Whether customer-data model training can be disabled |
Likely use cases described include sales training, customer-service conversations and management preparation. The product builds on Synthesia’s existing platform for AI-generated video.
Details such as whether the product is available on all plans or enterprise only, precise data-retention defaults, and the complete range of analytics available to managers should be confirmed directly with Synthesia before deployment.
What is AI roleplay coaching?
Traditional online training typically asks an employee to watch a video, read material or complete a quiz. The employee is a passive recipient of information.
AI roleplay allows the employee to enter a simulated scenario, speak or respond to an AI character, receive changing responses, practise handling resistance or uncertainty, repeat the session as many times as needed, receive feedback on their responses and compare their responses with a defined rubric.
A useful analogy
A training video explains the conversation. AI roleplay lets the employee rehearse it.
AI coaching systems typically combine several technologies: speech recognition to transcribe what the employee says, a language model to generate the avatar’s responses, synthetic voice to deliver those responses, an AI avatar to give them a visual form, scenario instructions that define the situation, a scoring rubric that defines expected behaviours, transcript analysis to compare the conversation with the rubric, and an analytics layer to report the results.
How an AI coaching session works
- 1The organisation chooses a scenario
- 2The organisation defines the role played by the avatar
- 3The organisation defines expected behaviours or outcomes
- 4The employee begins the simulated conversation
- 5The avatar responds dynamically based on the employee’s input
- 6The system records or transcribes the exchange
- 7The interaction is compared with the rubric
- 8Feedback or a score is generated
- 9The employee may repeat the exercise
- 10Managers may receive aggregated or individual analytics depending on the system configuration
Where the judgement lives
The avatar may look like the coach, but the rubric is where the organisation’s judgement is encoded.
Why businesses may find it attractive
AI roleplay coaching could potentially offer several practical benefits over conventional training:
- Practice available on demand, without scheduling
- Consistent scenarios delivered the same way each time
- Employees can repeat sessions without embarrassment
- Rare situations can be rehearsed that rarely arise in real work
- Multilingual delivery potentially available
- Scalable across large or distributed teams
- Remote and hybrid workers can access training without travel
- Faster creation of scenarios compared with live roleplay sessions
- Analytics showing common learning gaps across the organisation
- Measurable participation records
An important qualification
Scalability is a genuine advantage, but scaling an inaccurate or unfair assessment would also scale the problem.
These benefits are potential, not guaranteed. They depend on scenario quality, rubric accuracy, system reliability, employee trust, accessibility and how results are governed. None of Synthesia’s described benefits should be treated as proven outcomes without independent evaluation in the specific organisation.
Coaching versus assessment versus monitoring
The distinction between coaching, assessment and monitoring matters enormously for data protection, employment law, employee trust and governance requirements.
| Dimension | Coaching | Assessment | Risk difference |
|---|---|---|---|
| Purpose | Improve confidence and skill | Judge competence or suitability | Assessment requires stronger justification |
| Visibility | Results may remain private | Results are visible to management | Private results carry lower monitoring risk |
| Consequences | No employment consequence | May affect appraisal, promotion or discipline | Employment consequences trigger higher legal risk |
| Validation | Helpful feedback may be sufficient | Scoring must be demonstrably reliable and fair | Assessment-grade validation is a higher bar |
| Data retention | Short or no retention may be appropriate | Formal records may be retained and challenged | Longer retention creates greater privacy risk |
| Human review | Human support available when needed | Human review must be meaningful before decisions | Approval without genuine review is not oversight |
The key risk
The same software can move from low-risk coaching to high-risk assessment simply because the employer changes how the score is used.
The monitoring risk
A tool can be described as coaching while still functioning as monitoring if managers can inspect, compare or act upon the results.
Who defines a good answer?
AI does not independently discover the organisation’s ideal behaviour. The system depends entirely on human choices made during design: scenario wording, rubric criteria, examples, scoring weights, thresholds, model prompts, language expectations and organisational assumptions.
Rubric criteria might legitimately include asking open questions, acknowledging concerns, explaining a process clearly, avoiding prohibited statements or reaching an agreed next step. These can be reasonable targets for a training exercise.
Problems arise when criteria inadvertently reward one preferred speaking style, excessive confidence, specific vocabulary, speed over thoughtfulness, scripted responses, cultural conformity or agreement with management assumptions. A rubric designed by a homogeneous team may reflect that team’s communication preferences.
The rubric principle
An automated score may appear neutral while reproducing the assumptions of the people who designed the rubric.
What an AI score can and cannot prove
| An AI score may help indicate | An AI score does not automatically prove |
|---|---|
| Whether expected topics were mentioned | Real-world competence |
| Whether a process was followed | Empathy |
| Whether prohibited wording appeared | Honesty |
| Whether defined questions were asked | Confidence |
| Whether the conversation reached a required stage | Leadership ability |
| Areas for further practice | Emotional intelligence |
| Suitability for promotion | |
| Ability under pressure | |
| Future job performance | |
| Intent |
Simulated performance in a training environment may differ from real workplace performance. A scripted scenario creates artificial conditions that do not capture every factor present in a genuine customer interaction, a difficult management conversation or a live complaint.
The evidence principle
A score is evidence produced by a model — not a fact about the employee.
Privacy and workplace monitoring
An AI coaching session may process a substantial range of personal data, including: employee name and ID, voice recording, transcript, session duration, responses, score, feedback, number of repeated attempts, performance trends over time, manager comments, and device or usage data.
Before deployment, UK organisations should establish: the purpose for which data is collected, the lawful basis for processing under UK GDPR, whether collection is necessary and proportionate, how employees will be informed, how long data will be retained, who may access it, how it will be secured, what employee rights apply, and what data-sharing arrangements exist with the supplier and subprocessors.
Merely calling a system “training” does not remove data-protection responsibilities. Where processing involves personal data in a new or significant way, the ICO recommends completing a Data Protection Impact Assessment before deployment. A DPIA should be considered rather than assumed unnecessary.
A critical point for employees
Employees should not discover during an appraisal that a supposedly private practice session became part of their performance record.
Automated decision-making
UK GDPR includes provisions concerning automated processing that produces decisions with legal or similarly significant effects on individuals. Not every AI coaching score automatically falls within these provisions. The risk increases as scores become more consequential.
Risk increases substantially when coaching scores are used to influence: recruitment decisions, probation outcomes, promotion, pay, shift allocation, performance ratings, disciplinary action, dismissal, access to further training, or role suitability assessments.
Where automated processing contributes to consequential employment decisions, human involvement must be informed, independent, capable of challenging the score, capable of changing the outcome, and genuinely engaged with the decision rather than merely approving the AI’s conclusion.
The oversight test
A manager clicking “approve” does not create meaningful human oversight if the AI score is never questioned.
Bias, disability and communication style
Speech-recognition and language-scoring systems may perform differently across different accents, dialects, first languages, speech patterns and communication styles. This is not hypothetical: research has consistently shown that speech-recognition systems have historically performed less accurately on some accent groups.
Employees with speech impairments, stammers, hearing impairments, neurodivergent communication styles or significant anxiety may participate differently in a voice-based simulation. Cultural communication styles vary: directness, pacing, formality, deference and the use of silence are culturally variable, and a rubric calibrated to one set of preferences may disadvantage others.
Equality Act 2010 obligations remain relevant. Where a system puts employees with a protected characteristic at a disadvantage, businesses must consider whether the practice is justified, whether adjustments are available and whether alternative participation methods exist.
- Test recognition accuracy and scoring consistency across different accents, dialects and languages
- Identify false positives and false negatives
- Provide alternative participation methods for employees who cannot use voice-based interaction
- Define a reasonable-adjustment process before launch
- Establish an appeal route for disputed scores
- Review scoring patterns regularly to identify potential bias
The fairness principle
Training should help employees develop; it should not penalise them for communicating differently from the model’s preferred pattern.
Emotion and personality inference
Some AI systems claim to infer emotion, engagement, confidence, honesty, attitude, personality or stress from voice tone, word choice, response speed or facial movement. These claims should be treated with significant caution. The scientific basis for inferring psychological states from voice and expression patterns is contested. No coaching tool should be trusted to assess confidence, empathy, honesty or emotional intelligence from surface indicators unless this is supported by robust independent evidence specific to that tool.
Public documentation does not currently confirm whether Synthesia Roleplay Sessions uses facial or emotion analysis. If the system does not use such analysis, suppliers should say so clearly. Businesses should confirm this position and not assume it.
Do not purchase a conventional conversation-scoring tool and later expand it into emotion or personality assessment without a fresh legal, ethical and technical review.
When AI coaching becomes higher risk
| Risk level | Characteristics |
|---|---|
| Lower risk | Voluntary use · private practice · no manager access to individual results · no long-term retention · no employment consequences · limited personal data · employee can repeat freely |
| Moderate risk | Manager sees completion status · aggregate team analytics · standard feedback · short retention · no formal decisions |
| Higher risk | Individual scores visible to managers · employee comparison · mandatory participation · persistent performance history · scores used in appraisals · scores affect promotion or pay · no appeal · no reasonable adjustment |
| Very high risk | Automated recruitment rejection · disciplinary action · termination · emotion recognition · covert monitoring · biometric categorisation · sole reliance on model score · no accessible alternative · no meaningful human review |
Where risk actually lives
The risk comes less from the avatar itself and more from the consequences attached to its judgement.
Human oversight
Trained humans should be responsible for approving scenarios before deployment, reviewing rubrics for fairness, testing scoring accuracy, reviewing complaints, considering individual context, identifying employees who need reasonable adjustments, challenging unusual results, monitoring for potential bias, verifying significant model changes, deciding whether results may be used formally and retaining ultimate authority over employment decisions.
Employees should have the right to understand the score and how it was produced, access relevant data held about them, correct factual errors, ask for human review, challenge the outcome, request adjustments and report inappropriate or distressing scenarios.
The IT Club position
AI should support the coaching conversation, not remove the human conversation.
Could it replace human coaching?
Not completely. AI may be genuinely useful for repetition, rehearsal, basic feedback on procedural compliance, low-stakes scenarios, preparing before a human coaching session and reinforcing standard processes. These are valuable contributions to a training programme.
Human coaches remain important for providing context and nuanced judgement, offering emotional support, understanding organisational dynamics, handling complex interpersonal issues, addressing safeguarding or discrimination concerns, supporting leadership development, enabling confidential discussion and challenging whether the rubric itself is appropriate.
The strongest model
The strongest model may be AI for frequent practice and humans for judgement, context and development.
Questions to ask an AI coaching supplier
- 1What exactly is recorded during a session?
- 2Is audio retained after the session ends?
- 3Is video retained?
- 4Is a transcript created?
- 5Are facial expressions analysed?
- 6Is emotion inferred from voice or behaviour?
- 7Is voice tone or voice characteristic analysed?
- 8Is any biometric data processed?
- 9Who creates the scoring rubric?
- 10Can customers modify the rubric?
- 11How is scoring validated?
- 12Has scoring been independently tested?
- 13How is performance across accents and dialects tested?
- 14How are disabled users accommodated?
- 15What languages are supported?
- 16What accuracy differences exist between languages?
- 17Can individual employee scores be hidden from managers?
- 18Can employees use the system privately?
- 19What analytics are available to managers?
- 20Can managers compare workers individually?
- 21Can results be exported?
- 22Can the system integrate with HR or learning-management platforms?
- 23Is customer data used to train supplier models?
- 24Can model-training use be disabled?
- 25How long is session data retained?
- 26Where is data hosted?
- 27Which subprocessors are used?
- 28Can individual sessions be deleted on request?
- 29Can an employee obtain their session data?
- 30How are model updates communicated?
- 31Can scoring change after a model update?
- 32What security certifications are held?
- 33How are data incidents disclosed?
- 34How does the supplier support a DPIA?
- 35What contractual protection is provided against intellectual-property or data-protection claims?
Questions to ask internally
- Why are we introducing the system?
- Is the purpose training, assessment or both?
- Is participation voluntary or mandatory?
- Who can see individual results?
- Will scores affect employment decisions?
- Could results be used for another purpose later?
- How long will session data be retained?
- Have employees been consulted?
- Have employee representatives or unions been involved where appropriate?
- Has a DPIA been completed?
- Has an equality impact assessment been completed?
- What reasonable adjustments are available?
- Is there a non-AI training alternative?
- How can employees challenge a result they believe is inaccurate?
- Who owns the rubric and who may change it?
- Who reviews accuracy after model updates?
- How will success be defined and measured?
- Who has authority to pause or stop the system?
Safe pilot approach
- 1Define a narrow use case — choose a low-risk training scenario with no employment consequence
- 2Keep it developmental — do not connect pilot scores to appraisals, pay or disciplinary decisions
- 3Consult employees — explain the purpose and gather concerns before launch
- 4Complete a DPIA — assess the personal-data and monitoring risks properly
- 5Conduct an equality review — test different accents, communication styles and accessibility needs
- 6Limit data collection — collect only what is genuinely required
- 7Keep results private where possible — allow employees to practise without surveillance
- 8Use volunteers for early testing — avoid coercive participation
- 9Provide an alternative — do not make the AI tool the only route to training
- 10Test the rubric — use multiple human reviewers and varied scenarios
- 11Review outputs — check for false or unhelpful feedback before wider rollout
- 12Retain human coaching — give users access to a person
- 13Set a short retention period — delete pilot data unless genuinely required
- 14Review before expansion — do not automatically extend coaching into assessment
The pilot principle
Pilot the system as a learning tool before considering whether it should ever become an assessment tool.
Immediate actions for businesses
- 1Inventory AI workplace tools — identify every system used for training, recruitment, monitoring or assessment
- 2Define the purpose — write down whether each tool coaches, assesses or monitors
- 3Map the data — record what is collected, stored and shared
- 4Restrict manager visibility — do not expose individual practice results by default
- 5Review the rubric — check what behaviour is being rewarded and by whom
- 6Complete risk assessments — consider privacy, equality, employment and security
- 7Create an appeal route — employees must be able to challenge inaccurate results
- 8Provide alternatives — support accessibility and reasonable adjustments
- 9Test before formal use — do not connect unvalidated scores to HR decisions
- 10Review contracts — check training use, data retention, security and liability
- 11Consult employees — explain the system clearly before launch
- 12Assign ownership — name a business, HR and technical owner
The governance principle
Decide what the AI score is allowed to influence before collecting the first employee session.
Warning signs
Investigate further or pause deployment if:
- Employees are not clearly told that sessions are recorded
- A training tool has been introduced without employee consultation
- Managers can see every individual practice attempt
- Scores are used in appraisals without prior validation
- The supplier describes scoring as objective without supporting evidence
- No one can explain what the rubric rewards or how scores are calculated
- Employees cannot challenge a result they believe is wrong
- No non-AI training option exists
- Disabled employees cannot participate on equal terms
- Accents or communication styles produce visibly inconsistent transcripts
- Emotion or personality is inferred from voice or facial movement
- Session data is retained indefinitely
- Customer data trains the supplier’s model by default
- Scores change after a model update without explanation
- The tool is integrated into HR systems without a fresh impact assessment
- Managers use individual league tables to make comparisons
- Employees are practising sensitive conversations using real names or live cases
- Disciplinary action relies mainly on AI output
- The organisation cannot delete session data when requested
- No named human owns the coaching process
Practical business implications
Training can become more active
AI roleplay may bridge the gap between watching and practising. Passive training asks employees to absorb information; active practice requires them to apply it. That can improve retention and confidence, particularly for situations that do not arise often in real work.
Private practice can improve confidence
Employees may feel more comfortable making mistakes with a simulation than in front of a colleague, trainer or customer. The ability to repeat sessions without embarrassment, try different approaches and pause when needed may make practice more accessible.
Scoring creates governance risk
A feedback tool becomes more consequential when results are stored, retained over time or visible to managers. What begins as developmental feedback can quietly become a performance record if governance controls are not designed and maintained from the start.
Workplace data requires transparency
Employees must understand what is collected, why it is collected, who can see it and how it may be used. Transparency is not only a legal requirement; it also affects whether employees trust and engage with the system.
A rubric is not objective merely because software applies it
Automation makes a process faster and more consistent, but consistency alone does not mean fairness. If the rubric is poorly designed, consistently applying it consistently produces unfair results at scale.
Accessibility must be designed in
Voice-first systems may not work equally for every employee. Accessibility is not an add-on; it must be considered from the outset. Where a system cannot adequately accommodate an employee’s communication needs, an alternative must be provided.
Model updates can change results
The same employee response may be scored differently after a supplier updates the underlying model. This creates a consistency problem if scores are used in formal records: the criteria for assessment may shift without the organisation knowing.
The standardisation risk
AI can standardise a training exercise without necessarily making the judgement fair.
The IT Club View
AI roleplay is one of the more credible workplace applications of generative AI. There is genuine value in allowing employees to rehearse difficult customer interactions, management conversations, sales objections, complaints and unfamiliar situations without requiring another employee or trainer every time. Available on demand, repeatable and increasingly capable of dynamic response, AI practice sessions can usefully supplement a training programme.
The concern begins when the organisation starts treating practice data as performance evidence. A worker experimenting in a training environment should be able to make mistakes, try different approaches, repeat the exercise and receive useful feedback without wondering whether an imperfect attempt will appear in their next appraisal.
The IT Club View
A safe place to practise stops being safe when every mistake becomes management data.
IT Club is not arguing against AI coaching, learning analytics, automated feedback, roleplay simulations or measurable training. The case is for clear purpose, employee transparency, private practice, limited retention, fair rubrics, accessibility, meaningful human review and a firm separation between coaching and employment decisions.
The IT Club View
Use AI to give employees more opportunities to learn — not to create a new layer of invisible surveillance.
Assessing AI coaching for your team?
Use our AI Employee Coaching Assessment Checklist to review purpose, employee privacy, scoring fairness, accessibility, manager visibility, data retention and human oversight before starting a pilot.
The value of AI coaching
AI coaching has clear uses: rehearsal, repetition, on-demand practice, scenario consistency and support for employees who find live roleplay uncomfortable. These uses do not require monitoring, persistent scoring or manager visibility to be effective.
A broader point
The value of AI coaching depends not only on what the system teaches, but on who can see the results and how those results are used.
The central principle
An AI coach can be a useful rehearsal partner. It should not quietly become an unaccountable performance manager.
Related business questions
What is AI employee coaching?
AI employee coaching uses artificial intelligence to give employees a way to practise workplace conversations. An AI system — often presented as an avatar — plays a role such as a difficult customer, a colleague or a manager. The employee responds and receives feedback. The system can respond dynamically and repeat the exercise as many times as the employee wishes.
What are AI roleplay sessions?
AI roleplay sessions are simulated workplace conversations between an employee and an AI system. Unlike a training video that the employee watches, a roleplay session requires the employee to participate. The employee may speak, type or respond to the AI avatar, and the AI generates contextually appropriate replies. Synthesia’s Roleplay Sessions product is one example of this approach.
How does AI roleplay training work?
A roleplay system typically combines speech recognition, a language model, synthetic voice, an avatar, a scoring rubric and an analytics layer. The organisation designs a scenario and defines what good responses look like. The employee participates in the conversation. The system scores the interaction against the rubric and produces feedback. The employee can usually repeat the session.
Can an AI coach improve employee performance?
AI coaching may support skill development by providing repeatable practice and consistent feedback. Whether it improves real-world performance depends on the quality of the scenario, the accuracy of the rubric, the employee’s engagement and how the coaching connects to wider training. Supplier claims about learning outcomes should be treated as claims until independently evaluated.
Is AI coaching the same as workplace monitoring?
Not automatically — but it can become monitoring if managers can see individual results, if session data is retained, or if scores influence employment decisions. A tool described as coaching may function as monitoring depending on how it is configured. The distinction depends on visibility, retention and consequences rather than on the product’s marketing name.
Can employers record AI coaching sessions?
Employers can potentially record coaching sessions, but must comply with UK GDPR and data-protection law. This requires a lawful basis for processing, appropriate employee notice, a necessity and proportionality assessment, defined retention limits, security controls and a DPIA where appropriate. Recording sessions without clear purpose or notice may breach data-protection obligations.
Can managers see employee coaching scores?
Whether managers can see individual scores depends on the system configuration. Some platforms allow managers to see all employee results; others provide only aggregate data. Restricting individual score visibility during developmental coaching reduces the risk of the system functioning as unannounced monitoring. Businesses should define visibility settings before launch and explain them to employees.
Can AI scores be used in appraisals?
Technically, scores can be retained and used in appraisals, but this creates significant governance risk. An unvalidated score may be inaccurate, biased or inconsistent. Using it formally without validation, human review or employee awareness could create employment-law exposure, equality concerns and employee-relations problems. If scores are to be used formally, they require validation, transparency and an appeal route.
Does UK GDPR apply to employee coaching tools?
Yes. AI coaching tools that process employee personal data — including voice recordings, transcripts, scores and performance trends — are subject to UK GDPR and the Data Protection Act 2018. Employers must identify a lawful basis, provide privacy information, ensure data security, respect employee rights and consider a DPIA where the processing involves new or high-risk activity.
Is a DPIA required?
A Data Protection Impact Assessment should be considered whenever AI coaching tools process personal data in a new or significant way. The ICO recommends DPIAs for processing that is likely to result in high risk to individuals, including monitoring employees or processing biometric data. A DPIA is not always mandatory, but completing one is good practice for any system that records, transcribes and scores employee interactions.
What is automated decision-making?
Automated decision-making is a process in which a decision with legal or similarly significant effects is produced wholly or significantly through automated means, without meaningful human involvement. UK GDPR provides individuals with rights in relation to certain automated decisions. Whether specific AI coaching use falls within those provisions depends on how consequential the automated output is and how much genuine human judgement is involved.
Can AI make employment decisions?
AI should not make employment decisions without meaningful human review. Using an AI coaching score as the basis for promotion, pay, disciplinary action or dismissal without a human genuinely considering other evidence and exercising independent judgement creates significant legal and ethical risk. Human oversight must be more than a formality.
Are AI coaching scores objective?
No. AI coaching scores reflect the rubric designed by the organisation, which reflects the assumptions and preferences of the people who created it. The model that applies the rubric also introduces variation. A score that appears precise is still the output of a system designed by humans with specific assumptions about what good communication looks like.
Could accents affect AI scoring?
Yes, possibly. Speech-recognition systems have historically shown lower accuracy on some accent groups. If transcription accuracy varies, scoring accuracy may vary in parallel. Organisations should test recognition and scoring consistency across the accents and dialects spoken by their workforce before connecting the system to any consequential decision.
How should disabled employees be supported?
Disabled employees must be able to participate on fair terms, which may require reasonable adjustments including alternative participation methods, extended time, human assistance, modified scenarios, or a different form of training altogether. Employers have a legal duty under the Equality Act 2010 to make reasonable adjustments. This obligation applies to AI coaching systems as it applies to other workplace tools.
Should coaching results remain private?
For developmental coaching, keeping results private from managers is generally lower risk and more consistent with a genuine learning environment. Where scores are visible to managers, employees should be told this clearly. Private practice encourages experimentation, reduces anxiety and increases the likelihood that employees engage honestly with the training.
Can AI replace a human trainer?
Not completely. AI can support rehearsal, repetition and basic procedural feedback, but human trainers remain necessary for nuanced judgement, contextual understanding, emotional support, safeguarding, leadership development and challenging whether the training approach itself is appropriate. The strongest model uses AI to increase practice frequency and humans for quality and context.
How long should coaching data be retained?
Retention should be limited to what is genuinely necessary for the stated purpose. For developmental coaching with no employment consequence, short retention is appropriate. Data should not be retained indefinitely by default. Retention periods should be defined before deployment, documented, and applied automatically. Employees should be told how long data is kept and have the right to request deletion where applicable.
What should employers ask an AI coaching supplier?
Key questions include: what is recorded, whether audio and transcripts are retained, whether emotion or facial analysis is used, who designs the rubric, how scoring has been validated, how different accents and communication styles are handled, whether managers can see individual results, whether customer data trains the supplier’s model, where data is hosted, and what contractual protection is provided. A full list of 35 questions appears in the article.
How should an AI coaching pilot be run?
Start with a low-risk scenario, volunteer participants, no connection to appraisals or employment decisions, private results, short retention and a defined stopping point. Complete a DPIA and equality review before the pilot begins. Test the rubric with varied accents and communication styles. Collect employee feedback. Review before any expansion. Do not treat an untested pilot as a validated system.
Can employees challenge an AI score?
They should be able to. Any system that scores employees — even in a coaching context — should allow employees to understand how the score was produced, identify factual errors and request human review. Where scores are used in formal HR decisions, a meaningful appeal process is not optional. Employees also have rights under UK GDPR to access personal data held about them.
Does the EU AI Act apply to workplace coaching?
The EU AI Act may apply to certain workplace AI systems depending on where the system is developed, supplied and used. Systems used for making employment-related decisions may carry higher legal and governance requirements than systems used only for voluntary, private practice. Applicability depends on the specific use case, jurisdiction and the extent of employment consequences. UK organisations are not directly subject to the EU AI Act, but should understand its implications if they supply products or services into the EU or use systems supplied from there.
Administrator Technical Note
This section is intended for IT administrators, data-governance leads, HR technology teams, legal and compliance staff, and technical owners responsible for deploying AI coaching platforms.
Technical architecture
A typical AI coaching deployment involves multiple technical components, each of which may involve a different supplier, subprocessor, data location, retention period, security control, model version and update cycle. Components likely include: user identity and authentication, training platform, scenario engine, large language model, speech-to-text, text-to-speech, synthetic avatar, scoring rubric engine, analytics engine, transcript storage, reporting dashboard, learning-management-system integration, HR-system integration, audit logs and supplier APIs.
The governance risk
The avatar is only the visible layer; the governance risk sits across the complete data and scoring pipeline.
Data flow mapping
Document the following flow: 1. Employee authenticates. 2. Scenario is loaded. 3. Audio, text or video is captured. 4. Content is transmitted to the platform. 5. Speech may be transcribed by a third-party service. 6. The language model generates responses. 7. The conversation is scored against the rubric. 8. Feedback is produced. 9. Transcript and analytics may be stored. 10. Results may be shared with managers or other systems. 11. Data may be used for service improvement. 12. Data is deleted or retained according to policy.
For each stage, record: controller, processor, subprocessor, data category, purpose, lawful basis, location, encryption, retention, access controls and deletion method.
Data minimisation
- Prefer pseudonymous or anonymous pilot accounts
- Avoid video capture where not required
- Avoid biometric analysis
- Use transcript-only processing where sufficient
- Restrict manager access to individual attempts
- Use aggregate analytics by default
- Set short retention periods
- Provide employee-controlled deletion where possible
- Use synthetic scenarios, not real customer names or live HR cases
- Exclude sensitive HR content from training scenarios
- Avoid production HR integrations during pilots
Do not collect facial, vocal or behavioural data simply because the platform can.
Scoring-rubric governance
For every rubric, record: purpose, owner, scenario, expected behaviour, prohibited behaviour, weighting, pass threshold, evidence base, training examples, supported languages, tested groups, accessibility considerations, reviewer, approval date, model version, last validation date, complaints received and change history.
Require version control, change approval, documented testing, human review and periodic revalidation as standard practice.
Validation testing
Test using: different accents, dialects, first-language speakers, non-native speakers, speech impairments, different pacing, concise responses, detailed responses, culturally varied communication approaches, assistive technologies, background noise, interrupted sessions and ambiguous scenarios.
Measure: transcription accuracy, score consistency, inter-rater agreement with human assessors, false positives, false negatives, repeatability, explanation quality, appeal outcomes and performance drift over time.
Do not use employee deployment as the first real validation test.
Security controls
Recommended controls: single sign-on, multi-factor authentication, role-based access, least-privilege principles, encryption in transit and at rest, restricted administrator access, audit logging, retention controls, export controls, secure deletion, tenant separation between customers, supplier incident notification, vulnerability management evidence, penetration testing evidence, API security review, integration access review and data-loss prevention where appropriate.
Protect particularly: transcripts, recordings, scores, manager comments, sensitive scenario content and authentication data.
HR and LMS integrations
Integration with HR information systems, performance platforms, recruitment systems, learning-management systems and CRM platforms can change the risk profile significantly.
A training result should not flow into a formal employee record merely because an integration makes it technically convenient.
Before integration, review: purpose, data fields, direction of transfer, permissions, retention, visibility controls, automated actions triggered, employee impact, rollback capability, deletion behaviour and auditability.
Model and scoring drift
Scoring behaviour may change due to: model updates, speech-recognition updates, rubric changes, prompt changes, supplier configuration changes, language-model replacement, safety-policy changes or analytics changes.
A recurring test should compare: standard benchmark conversations, previous scores, current scores, explanation quality, subgroup performance and error rates. Run drift testing after every significant model update.
Employee data inventory
Recommended fields: platform, supplier, business owner, HR owner, technical owner, purpose, users, voluntary or mandatory, data captured, recording status, transcript status, score status, manager visibility settings, HR decision use, lawful basis, DPIA status, equality assessment status, retention period, data location, subprocessors, integrations, current model version, rubric version, appeal process, last validation test, incidents, next review date and decision.
Incident response
Potential incidents include: inaccurate scores, exposed recordings, unauthorised manager access, inappropriate scenario output, discriminatory feedback, incorrect transcript, accidental use of customer data, excessive retention, score changes following a model update, data exported to the wrong system, employee complaint, and supplier breach.
- 1Pause affected use
- 2Preserve relevant evidence
- 3Restrict access
- 4Identify affected employees
- 5Review model and rubric versions at time of incident
- 6Assess data-protection impact
- 7Correct employee records where necessary
- 8Notify the supplier
- 9Notify affected employees where required
- 10Consider regulatory reporting obligations
- 11Retest before re-enabling the system
- 12Document corrective action and review schedule
Operational Heartbeat
AI coaching risk changes as models are updated, scenarios are added, rubrics change, managers request more analytics, integrations are enabled, employee participation expands, retention grows, supplier terms change, subprocessors change, new languages are added, scoring methods change, legal guidance develops, complaints arise and the tool moves from coaching into assessment.
A recurring review should check: purpose, employee transparency, active scenarios, rubric versions, current model versions, accuracy testing, equality testing, accessibility arrangements, manager permission settings, retention, customer-data training settings, integrations, complaints received, appeals, supplier incidents, contract changes, DPIA status, corrective actions, named owners and next review date.
Operational Heartbeat
AI coaching needs an operational heartbeat: purposes, rubrics, model versions, permissions, retention and employment consequences should be reviewed rather than allowed to expand silently.
Plain-English Takeaway
AI roleplay tools can give employees a convenient and private way to practise difficult workplace conversations. The risk increases when sessions are recorded, scored, compared or used in appraisals and employment decisions. Businesses should explain the purpose clearly, minimise the data collected, test the scoring for fairness, provide human review and keep developmental coaching separate from formal performance management.
Need the practical steps?
A short, instruction-led version of this topic is available in the Knowledge Centre.
View the Knowledge Centre GuideRelated Articles
Why Do AI Companies Want Old Books?
Reports suggest that AI developers and specialist data suppliers are buying large quantities of second-hand books — some to be scanned destructively and then discarded. The reason is not nostalgia. It is data quality, copyright and the growing value of verifiably human-created knowledge.
Read articleDid an AI Really Escape and Launch a Cyberattack?
A reported AI security incident has attracted dramatic headlines. But the most important lesson is not about a machine gaining consciousness. It is about what happens when a capable system is given powerful tools, broad access and insufficient containment.
Read articleHow to Find Every Photo on Your Windows PC
Photographs often become scattered across a Windows PC — saved in Pictures, Downloads, Desktop, OneDrive, project folders and places you have long forgotten. File Explorer can search across the entire computer and display image files from many locations in one results view. Here is how to do it safely.
Read article