I spoke with Anke Paulick, President of the ICF Germany Chapter, on the Speexx Exchange podcast about AI coaching governance and the AI coaching audit recently conducted by the ICF. I asked her what the first independent, ICF-aligned audit of an AI coaching platform means for enterprise HR and L&D leaders. Anke’s answer reflected the International Coaching Federation’s role: coaching quality has to move from debate to disciplined practice, from curiosity to clear criteria, and from pilots to reproducible professional review.

The most useful moment in the conversation was the shift from claims to audit work. Marketing language can be very smooth in AI coaching: aligned with standards, built with experts, safe by design, human-centered, scalable, personalized. The audit language Anke brought into the discussion was much harder to decorate. A 69-item checklist. Nine compliance categories. Twelve documented test cases with transcripts. Stress, cultural sensitivity, crisis, and manipulation. That vocabulary changes the conversation entirely. That audit lens is the difference between a vendor saying “trust us” and a buyer asking “show us.” AI coaching standards do not remove judgment from the buying process. They make judgment more precise.

Content

Standards protect the profession and the buyer at the same time

The ICF AI Coaching Framework is built around six domains. Foundation covers ethics, mindset, and professional identity. Co-Creating covers trust, safety, and the coaching agreement. Communicating covers active listening and powerful questioning. Cultivating covers learning, growth, and client accountability. Assurance and Testing covers feedback, bias monitoring, and validation with diverse users. Technical Factors covers security, privacy, resilience, and accessibility.

That structure is useful because it refuses a false choice. AI coaching is not simply a technology product, and it is not simply a coaching conversation delivered through a new channel. It is both. The ICF AI Coaching Framework and Standards gives providers, buyers, and coaches a shared language for evaluating the full system.

In the podcast, Anke made the professional consequence clear. If the coaching profession does not define guardrails, others will define them too low. That warning applies to market claims as much as ethics. A vendor can claim “ICF-aligned” without evidence. A buyer can accept a demo as proof. A procurement team can reduce coaching to a price-per-user calculation. Standards interrupt that drift.

The result is not bureaucracy for its own sake. Standards help buyers protect value. HR and L&D leaders want democratization of coaching, more access for employees outside the executive population, personalized support, and more coaching cases for the same budget. Those goals are sensible only if the service being scaled still behaves like coaching. Standards protect the economics by protecting the substance.

The audit mindset changes what buyers ask for

The audit Anke discussed followed three broad steps: self-assessment, expert review, and calibration. Speexx documented the platform against a 69-item checklist across 9 compliance categories. The ICF Germany team reviewed documents and tested real coaching scenarios. The calibration step included 12 documented test cases with transcripts, including stress, cultural sensitivity, crisis, and manipulation.

That structure is important because it combines documentation with behaviour. Documentation shows what the provider intends. Hands-on testing shows what the system does. Calibration creates a disciplined way to compare findings, refine judgments, and avoid treating isolated impressions as evidence. This is the difference between a compliance review and a professional coaching audit.

A normal software review can answer security, architecture, availability, and data-processing questions. Those questions remain necessary. They are not enough. AI coaching also requires assessment by people who understand coaching: boundaries, contracting, reflective space, non-directive questioning, values, accountability, and handover when the conversation has left coaching. The audit mindset brings coaching professionalism into the assessment of AI behaviour.

The ICF practical guide supports this broader view by helping stakeholders think through how AI and coaching should integrate. The guide is useful for buyers because it treats the standard as something to apply in real systems, not a document to cite once in a sales conversation.

Webinar Replay AI coaching standards and audit evidence

Listen to the Speexx Exchange Podcast with ICF president Anke Paulick and her view on quality standards in AI coaching and how to develop a real world check list.

Watch Now

Testing real scenarios is where quality becomes visible

Anke’s point about real scenarios was one of the most practical parts of the conversation. A system may perform well in a controlled demo and still fail when the user becomes ambiguous, emotional, culturally specific, manipulative, or highly directive. Coaching quality often appears in how the system refuses a wrong move, not in how well it answers an easy prompt.

The test cases Anke described included stress, cultural sensitivity, crisis, and manipulation. The script notes also refer to feedback conversations, conflict resolution, distress simulation, commercial steering, and crisis escalation. These are exactly the situations enterprise users bring to coaching. A manager preparing for difficult feedback is not testing a generic conversation. A leader in a cross-border team may be dealing with language nuance, power distance, regional norms, and personal credibility. An employee in distress may use coaching language while needing human support.

A serious audit asks whether the system holds the coaching stance under those conditions. Does it keep the user’s agency? Does it avoid giving inappropriate advice? Does it recognize distress signals? Does it hand over when needed? Does it avoid reinforcing bias? Does it stay transparent about its own limits? These questions are not rhetorical. They are product requirements.

This connects directly to the European AI governance context. The EU AI Act frames AI around risk, trust, and accountability, while the Commission’s guidelines for high-risk AI systems push providers and deployers toward more precise classification and evidence. AI coaching buyers may not classify every use case the same way, but they should apply the same habit of operational scrutiny.

Audit findings are useful only when they change the system

One of the reasons I found the audit conversation valuable is that it did not present audit as a ceremonial moment. Anke’s view was closer to professional learning. An audit does not end the conversation. It begins a disciplined one. That is an important distinction for providers and buyers.

In the session, crisis detection and escalation became the clearest example of why a serious audit matters. This is the point where AI coaching reaches its most sensitive boundary: the moment a conversation is no longer safely inside coaching and requires human support. The audit gave Speexx hard findings, and the value was in what happened next. Speexx responded quickly, took the recommendations seriously, strengthened the distress detection and escalation approach, and moved the findings into positive compliance. For me, that is the real value of an audit. It is not a badge. It is a disciplined way to improve quality with the industry, under the standards of the coaching profession.

The operational requirement is clear. An AI coaching system must detect distress signals across severity levels, have a defined escalation protocol, trigger handover when crisis signals appear, and document and test the policy. It must never coach suicidal behaviour independently, manage complex emotional crises without human handover, decide alone whether a user is genuinely at risk, or replace organizational safeguarding obligations.

This is where standards become more than an external badge. A provider has to be able to update policy, product behaviour, documentation, and testing. Buyers should look for responsiveness because AI coaching will continue to change. The strongest provider is not the one that claims perfection. It is the one that can show serious improvement against a professional standard.

Fosway, suite thinking, and the enterprise reality of audits

Fosway’s 2026 recognition of Speexx as a Core Leader in the Fosway 9-Grid for Digital Learning provides useful enterprise context. Speexx is not operating only as a narrow language or coaching provider. Speexx is a digital learning suite for communication capability, which means AI coaching sits inside a wider system of human expertise, learning infrastructure, measurement, and enterprise integration.

For audits, enterprise quality is rarely confined to the product screen. It includes onboarding, data protection, accessibility, support, integration, reporting, service delivery, change management, and governance. standards and certifications support that wider buyer confidence. Awards and certifications provide additional market evidence, but audit scrutiny remains different because it tests specific claims about AI coaching behaviour.

Speexx AI Coaching and Speexx Coaching™ belong in this suite logic. The platform supports human coaching, team coaching, group coaching, and AI coaching, while the broader Speexx environment connects coaching with language development, mentoring, intercultural programs, and capability intelligence. AI coaching becomes one delivery model inside a governed communication capability layer.

This is also where enterprise buyers need to resist simplistic comparisons. A standalone AI coaching tool may look fast and inexpensive. A suite-level solution has to answer harder questions about governance, integration, data, quality, and global rollout. The buying decision depends on whether the organization wants a conversational feature or a governed development capability.

Standards are the beginning of market maturity

The market needs standards because AI coaching will scale faster than the professional conversation around it unless buyers demand better evidence. Demos will improve. Interfaces will become more natural. Models will become more capable. None of that removes the need for coaching logic, tested safeguards, transparent limits, and clear accountability.

Anke’s ICF perspective keeps the profession involved where it belongs. The coaching profession should not admire innovation from a distance or reject it on instinct. It needs to ask harder questions and evaluate anything called coaching through the lens of coaching professionalism. That was the strongest professional message in the podcast.

For HR and L&D leaders, the implication is practical. Ask for the standard. Ask for the audit logic. Ask for the test cases. Ask what changed after the audit. Ask who is accountable. Responsible AI coaching will not be defined by the provider with the best phrase on the website. It will be defined by the provider that can withstand evidence.

Review the Speexx standards and certifications approach, then explore how Speexx AI Coaching fits into governed Speexx Coaching™ for enterprise L&D.
Book a demo
Request your Speexx Demo now!