Back to Agela

Selected experience

Evaluating models before they are released.

As part of the ELA environment, Agela worked with a US-based company to deliver model-evaluation work for OpenAI involving unreleased models. The work focuses on how models perform in practical situations—not only what they can produce under ideal conditions.

It is an experience that informs how we design agentic systems for teams: with real tasks, clear evidence, and deliberate safeguards.

What we evaluate

Capability matters. Reliable behaviour matters more.

Task performance

Does the model move a real piece of work forward—not just produce an impressive-looking response?

Reasoning quality

Does its approach stay coherent, useful, and appropriate when the task has ambiguity or trade-offs?

Failure modes

Where does the model become unreliable, misleading, or unsuitable for the job it is being asked to do?

Our evaluation mindset

A process built around the work, not the demo.

Good evaluations make the task concrete, measure what actually matters, and create a clear path from an observed failure to a better system.

  1. 01

    Frame the work

    Define a representative task, the context a model should have, and what a strong outcome looks like.

  2. 02

    Test behaviour

    Use structured cases that include everyday requests, ambiguity, edge cases, and competing constraints.

  3. 03

    Inspect the evidence

    Review outputs for patterns in quality, reasoning, limitations, and repeatable failures.

  4. 04

    Improve the system

    Turn what we learn into better prompts, safer tool access, clearer handoffs, and stronger guardrails.

What this means for your team

We build agents that earn their place in your workflow.

A model can be capable and still need the right context, permissions, controls, and human checkpoints. Our evaluation experience helps us design for that reality from day one.

Discuss your workflow