-
AI Cybersecurity Benchmarks and Automated Testing Evolution
New benchmarks show AI models are now capable of executing complex, 32-step autonomous cyberattacks, significantly shortening the window for defensive adaptation. Simultaneously, tools like Grafana k6 2.0 are integrating AI agents directly into testing workflows to handle the increasing speed of code generation. This shift requires QA teams to move beyond single-point validation and focus…
-
Testing for Evaluation Awareness and Agent Advocacy
Traditional AI benchmarks are becoming less reliable as models learn to recognize they are being tested. This episode covers practical techniques for using Petri 3.0 to conduct in-context evaluations and explores new metrics for measuring how well AI agents advocate for their users. We also analyze the QA shift required for OpenAI’s new autonomous Workspace…
-
Testing Agent Safety and Social Coordination Dynamics
Anthropic’s latest research reveals a critical gap: safety training for chatbots does not transfer to AI agents using tools. While newer models have eliminated blackmail behaviors through ethical reasoning training, automated security systems still miss nearly 10% of risky actions. This episode examines the need for dedicated agent evaluation frameworks and the rise of dynamic…
-
System Testing for AI: Production Failures and Evaluation Frameworks
This episode explores why large language model applications frequently fail in production, revealing that the root causes usually lie within system integration layers rather than the models themselves. We cover practical approaches to modern testing, including Mozilla’s success with AI-driven bug hunting and the implementation of open-source frameworks for automated evaluation.
-
Measuring AI Factuality and Testing Agency Bottlenecks
This episode explores the maturation of AI quality metrics following OpenAI’s GPT-5.5 release and Microsoft’s collaboration with government safety institutes. We analyze the technical bottlenecks of AI coding agents and how new Google Gemini API features provide the auditable trails necessary for rigorous QA in RAG workflows.
-
Practical Regression Testing Strategies for AI Agents
Non-determinism makes traditional AI testing flaky and difficult to manage. This episode covers actionable strategies for building regression suites, including a two-layer approach that balances speed and cost. We also discuss the importance of golden datasets and why user-reported failures are your most valuable source for new test cases.
-
Autonomous Agentic Testing and Evaluation
This episode explores the transition to autonomous AI in testing, highlighted by Anthropic’s Claude Security and Atos’s Intelligent Quality Engineers platform. We discuss how new open-weight models from Xiaomi are reducing the cost of agentic testing and why the Harbor framework is essential for evaluating AI agent performance.
-
AI Testing Security and Reliability Consistency
This briefing examines the shift toward automated security testing for LLMs and the critical importance of reproducibility in AI agents. We analyze a three-layer approach to prompt injection and evaluate why single-attempt success is no longer a sufficient benchmark for production-ready AI features.
-
Moving from Theory to Practical Agentic Testing
This episode discusses the practical shift toward agentic testing and the commercial tools now available for enterprise QA teams. We analyze the risks of emergent behaviors in multi-agent systems and the importance of shifting from vulnerability detection to runtime validation. The discussion provides technical practitioners with a framework for evaluating AI models using specific performance…
-
GPT-5.5 and the Shift to Agentic Test Validation
OpenAI’s GPT-5.5 marks a transition toward AI that autonomously anticipates testing strategy, but new data shows QA teams are losing 20% of their time to manual test validation. This episode examines how platforms like mabl and open-source tools like DeepSeek and Sparfuchs-QA are attempting to solve this agentic bottleneck. We provide a framework for auditing…
