Evaluating Voice AI Models in Hours for Consumer Call Screening

Client:

Consumer privacy and security company

Time to First Value:

4 weeks

Users Enabled:

Engineering and product teams

Our Problem

Problem

An AI call-screening product has to make a consequential distinction: let legitimate callers reach the customer while screening spam and scams. It also needs to keep conversations responsive and operating costs manageable.

‍

After observing changes in product behavior following a model upgrade, the company needed clearer evidence for choosing between model and prompt configurations. Manual and script-based tests were already in place, but the team lacked a repeatable way to compare changes across representative call scenarios.

‍

Overall accuracy could only tell part of the story. The team needed to understand which calls were handled incorrectly, how conversational behavior changed, and what each configuration would cost.

Our Solution

Solution

Eliza worked alongside the company’s engineering and product teams, with support from OpenAI, to establish a repeatable testing workflow around OpenAI’s realtime voice models.

The work focused on four areas:

  • Representative call scenarios: Defined test cases, routing expectations, failure modes, and quality criteria.
  • Reusable audio tests: Built synthetic audio assets with varied background conditions to test behavior across different call environments.
  • Model and prompt comparisons: Ran experiments through the company’s evaluation infrastructure and compared routing decisions, conversational quality, and estimated cost.
  • Human-reviewed scoring: Reviewed automated judgments and investigated failures to make evaluation results more useful for engineering decisions.

The company owned its native evaluation infrastructure. Eliza contributed experiment design, synthetic data, diagnostic analysis, and recommendations, then handed over documentation so the team could continue testing independently.

Our Results

Results

  • New models evaluated in hours, using a repeatable testing workflow.
  • Consistent comparisons across model and prompt configurations using common test scenarios.
  • Visibility into specific failure patterns and quality-cost tradeoffs that overall accuracy could miss.
  • An evaluation capability the customer could operate independently, supported by reusable test assets and handoff documentation.

The experiments identified a promising OpenAI realtime configuration for further optimization. More importantly, the team gained a practical way to assess future changes: run the tests, inspect failures, compare tradeoffs, and decide what to test next.

The engagement delivered the evaluation capability and evidence-backed recommendations. Production monitoring and rollout remained subsequent steps owned by the customer.