Archive

  • Building Local AI Testing Agents and Governance

    This episode examines the practical implementation of open-source models and enterprise governance in AI testing. We analyze how Google’s Gemma 4 and Microsoft’s Agent Evaluation platform allow testers to build private, auditable workflows. The discussion highlights why human testers must remain the primary drivers of exploratory reasoning despite advances in autonomous agent architecture.


  • GPT-5.4 Surpasses Human Benchmarks and SmartBear Launches Major AI Update

    This episode explores the technical milestone of GPT-5.4 surpassing human benchmarks for computer use and what it means for autonomous test execution. We cover SmartBear’s platform-wide AI rollout and the emergence of sovereign AI testing platforms. The discussion focuses on practical strategies for multi-model validation to eliminate AI hallucinations in QA workflows.


  • Multi-Model Verification and the Reality of Agentic Reasoning

    This episode examines how multi-model verification and formal mathematical proofs are setting new standards for AI reliability in software testing. We discuss the significant reasoning failures of frontier models in the ARC-AGI-3 benchmark and the practical application of Google’s new Java-based agent framework. This session provides a factual look at moving from probabilistic testing to…


  • Practical AI Testing: Red Teaming, Chatbot Scenarios, and Multi-Source Test Design

    This episode examines three practical shifts in AI testing: automated security red teaming, goal-based chatbot scenario validation, and multi-source AI test generation. It explains how Promptfoo, DeepEval, and Rasa fit different testing needs while outlining how the QA role is moving from manual test writing toward review and refinement of generated suites.


  • Testing AI Agents, Securing LiteLLM, and Reducing UI Test Brittleness

    Today’s episode explores how AI testing is evolving across workflow validation, dependency security, and UI automation maintenance. It covers Solo.io’s agentevals for full agent trajectory testing, the LiteLLM compromise as a direct attack on AI tooling, Rapise 9.0’s self-healing locator repair, and Galtea’s market signal that enterprise AI testing now requires dedicated investment.


  • Testing Agentic AI Across Desktop, Coverage, and Multi-Agent Systems

    This episode explores how AI agents are expanding the scope of software testing beyond individual applications into full desktop and multi-agent environments. It reviews new data showing higher test coverage from specialized agents and examines tools designed for testing non-deterministic systems. It also outlines practical implications for permissions, data handling, and system-level validation.


  • From Vibe Testing to Eval-Driven AI Testing

    This episode examines how AI testing is evolving from subjective checks to structured evaluation. It covers frameworks for testing multi-turn conversations, introduces practical AI security testing techniques, and explains how to apply a testing pyramid to agent-based systems.


  • Testing AI Beyond Lab Conditions

    This episode examines how AI testing is evolving beyond controlled environments. It highlights practical techniques for testing voice systems with real-world audio, evaluating performance under load, and analyzing failures in multi-step AI workflows. The focus is on making testing reflect how systems actually behave in production.


  • AI Testing Signals: Autonomous QA, Dev Tool Wars, and Security Standards

    This episode explores four major signals shaping AI testing: the rise of autonomous QA platforms, the expansion of AI into full developer workflows, the introduction of standardized AI security testing, and emerging federal evaluation standards. Together, they point to a shift from testing code to validating entire AI-driven systems.


  • Building Practical AI Testing: Evals, Chatbots, and Playwright MCP

    This episode focuses on practical techniques for testing AI systems, including building an LLM evaluation pipeline and automating chatbot testing with simulated conversations. It also explains why traditional testing methods fail for AI and how tools like Playwright MCP introduce new approaches to browser automation.