
Conversational AI is fundamentally different from traditional software. You cannot rely solely on standard pass/fail UI tests when a system generates dynamic, unpredictable, multi-turn responses.
To build trust in an AI system, you have to test both the underlying engineering and the human experience. I provide specialist quality engineering for chatbots and conversational AI, ensuring your system is technically robust, logically sound, and genuinely inclusive to talk to.
The Approach#
I blend heavy-duty automated API testing with deep exploratory UX testing to evaluate AI systems from the data layer up to the conversational interface.
My focus is on outcomes: ensuring the NLP (Natural Language Processing) accurately understands intent, the context is retained across long conversations, and the chatbot’s personality does not cause cognitive friction for the user.
1. Automation & Infrastructure#
AI requires a rock-solid data layer. I design, build, and execute automated test frameworks specifically for conversational systems.
- API & Middleware Testing: Automating tests for API middleware using tools like Python, Playwright, Mocha, and Chai.
- System Reliability: Scripting the heavy lifting—capturing and refreshing Bearer tokens, managing configuration, and continuously validating API health and response performance.
- Scale & Scenarios: Developing native frameworks to test API middleware against thousands of conversational scenarios.
- Monitoring & Reporting: Setting up real-time API monitoring (via cURL and shell automation) and visual dashboards like Allure so your team always knows the system’s health.
2. Conversational Logic & Accuracy#
A chatbot is only as good as its memory and comprehension. I build structured and dynamic test suites to rigorously assess the AI’s “brain”:
- NLP Accuracy: Verifying that the system correctly parses complex or ambiguous user intents.
- Context Retention: Ensuring the AI remembers facts and context across long, multi-turn conversations.
- Evidence Summarisation: Validating how accurately the system summarises information from large knowledge models.
3. Exploratory & UX Testing (The Human Factor)#
An AI can pass every API check and still be awful to interact with. I conduct extensive exploratory testing focused on the actual human experience.
- Neurodivergent Accessibility: Evaluating the cognitive load of the AI’s responses to ensure they are accessible and do not overwhelm users.
- Tone & Empathy: Testing the “chatbot personality” to ensure responses remain empathetic, appropriate, and inclusive under stress or unexpected inputs.
Experience in AI Delivery#
I have delivered quality engineering for custom AI and chatbot systems across major compliance and telecom sectors. Highlights include:
- Acorn Compliance: Built automated and exploratory frameworks for a conversational AI steering users through complex DTAC compliance processes. Tested for NLP accuracy, multi-turn logic, and evidence summarisation across large DTAC knowledge models, while heavily testing the chatbot’s empathy and tone for neurodivergent accessibility.
- Tele2: Delivered quality engineering for a custom AI chatbot built on the IBM Watson service, developing a native Python 3 framework to test API middleware across thousands of user scenarios.
Are your AI systems release-ready?#
If your team is building a conversational AI product and needs a testing strategy that goes beyond basic unit tests to tackle real-world unpredictability, we should talk.
