Curlscape logo

Evaluating AI agents for Engineering Workflows

Evaluating AI agents for Engineering Workflows

TL; DR

This is the second post in the series “Agentic AI”. The first post described how an agentic engineering workflow differs from a scripted one and ended with four questions about how to trust its output. This one tries to answer them with data.

AI agents can now run an engineering simulation from start to finish. They set up the case, build the mesh, run the solver and report the results. For using them in production, the hard question is whether they can be trusted. There are multiple open evaluation suites out there where one can run the evaluation. One such example is HWE-bench - a public benchmark that grades AI agents on real OpenFOAM simulation tasks.

HWE-bench is an evolving dataset of various engineering simulations. At the time of us running the evaluation, this repo contained 22 CFD cases. It covers a range of simulations - basic ones like Lid driven cavity flow and NACA 0012 subsonic to more advanced ones like oblique shock and 3D Rayleigh-Bernard Convection.

This post covers details on evaluation harness (the details on evaluation process), the scores and some comments about the evaluation.

Note:

Note that, the HWE-bench benchmark results are all single-pass ones, whereas we have taken an agentic approach. While this is not an apples-to-apples comparison, the comparison itself is to demonstrate how language models can comprehend simulation tasks.

Setup Overview

The system under test is a multi-agent pipeline for simulation work. It takes a written request in plain English and runs it through eighteen agents in five phases (planning, pre-processing, setup, solution and post-processing), driving OpenFOAM. Reviewer agents check the output of the meshing, setup, solution and post-processing stages and can send the run back for another attempt. Every run used Claude Opus 4.6.

The benchmark is HWE-bench, published by SVD Lab. In each case we give the agent the following:

  • A REPL environment where it can execute commands
  • A working simulation solver (in this case it was OpenFOAM. But in principal, this can be any other solver as well)
  • An engineering problem written in plain English stored in instruction.md in each case.

For every case we pasted the benchmark's instruction.md verbatim into our multi-agent system. We ran 19 of the OpenFOAM cases and all 19 produced valid scores.

One caveat to call out here, we run OpenFOAM inside our own container as opposed to HWE-bench which uses their own system called harbour.

The section below talks about how the grading is done.

How HWE-bench grades a run

HWE-bench evaluates the agents using a deterministic script based approach. This makes it very easy to use since all we need to do is to run the script once the simulation is completed. More concretely, we follow the steps:

Agent reads the instructions from instruction.md file

LLM based agent (or the intelligence layer - in HWE-bench’s case it is a single LLM) then creates a geometry, meshes it, sets up the case, solves it

The reference solution script (each simulation has a reference script, which contains the right way to run it) runs and produced reference output

We use the script (in some cases, a command) that reads simulation output files along with the reference output and comes up with a score. This score is referred to as KPI, and they are named. Some KPIs measure fields outputs like velocity at the centre of the domain and others measure the mesh count etc.

A single simulation can have a combination of KPIs - some targeting meshing, some physics setup and others outputs themselves.

The benchmark's worked example looks like this:

{

"u_centerline_y0p5": {

"value": -0.2058,

"source": {

"kind": "file_extract",

"path": "/root/case/postProcessing/sets/200/centerline_U.csv",

"extract": "awk -F',' '$1==0.5 {print $2}'"

}

}

}

There's no LLM judge anywhere in the scoring. That takes away the headache of scoring the scorer.

Each reported quantity (a KPI) is scored as

source_verified × physics_pass × T_decay.

What each of these mean:

  • Source verified - (either 0 or 1) - checks whether a reported value can be reproduced from the stated source.
  • Physics Pass - (either 0 or 1) - Checks if the output is in the range allowed by physics
  • T decay - (continuous value between 0 and 1) - measures agreement with the reference value. HWE-Bench has a minimum and a maximum value specified, based on which, we calculate this score linearly. Below is one example of how T Decay value is used. This value is for a case of backward facing step simulation where the value of T decay depends on the re-attachment length. The reported value is around 5 mm, any value ±1mm around that is acceptable but beyond that, it penalises the final KPI value.
Worked example of HWE-bench scoring, showing how T_decay changes with reattachment length and how KPI groups are weighted in the final score.

Further, we add a weighting on each step to score the agent for getting the intermediate steps correctly as well.

Some nice attributes of KPIs and scoring process:

It gives marks for intermediate steps: The grader gives every failed KPI two labels, one for how far the simulation got and one for how close the value is to the source.

Labelling the source of error: A KPI can score zero because the physics is wrong or because a correct value couldn't be traced, and the labels tell you which one to fix.

The results

The single-agent column is SVD Lab's published Claude Opus 4.6 run from the v0.2 and v0.3 results, which makes it the same model running as one agent with a shell.

Discussion

General Observations

Off all the runs, 8 are within 0.07 of HWE-bench

We faced a bunch of issues in how we ran the simulations. We fixed them at the time of evaluation step

Case-specific Discussion

Lid-driven cavity at Re = 400

  • The reference value is wrong. For this simulation, grader assumes −0.16914 to be the average velocity at y=0.5, and cites Table I in Ghia, Ghia and Shin (1982).
  • If you check the table, you will notice the value being −0.11477 at y = 0.5.
  • Our agent’s results agree with Ghia better than their OpenFOAM results. The benchmark's own oracle gives −0.1105 and our solver evaluated −0.11718.
  • It caps every correct answer. Against −0.16914, no correct submission can score above about 0.68 on accuracy. Our validation reviewer flagged the discrepancy during the run, prompted by the brief's reference to Ghia's table.

Dam break

  • The brief and the reference geometry disagree. The brief describes a plain rectangular tank. The reference geometry copies OpenFOAM's damBreak tutorial verbatim, and that tutorial has a 0.024 × 0.048 m obstacle on the floor that the brief never mentions.
  • The obstacle changes the flow. Water hitting it throws up a fast vertical jet, which a plain tank doesn't produce.
  • Our score reflects that gap. Our run followed the brief and scored 0.926, and every lost point is in the one KPI the obstacle would affect. We haven't re-run with the obstacle, so the link is strongly suggested but not yet proven.
  • The KPI itself is inconsistent. It is defined four different ways across three files in the case directory.

Four cases graded against a coarse reference solve

In the cases specified in the table below, the scorer relies on using their own reference run in order to score the simulation results. This is, despite the fact that there is literature available to quote.

Each case's calibration note admits that solve disagrees with the experiment, so a run that matches the experiment loses points, even if the intent here is to make sure simulation is more accurate.

  • NACA 4412: our lift and drag were close to the experiment, and both scored zero for being too far from the oracle. That caps our score at 0.616, although it's still well ahead of the single agent's 0.250.
  • Backward-facing step: reporting the experiment's own reattachment length (6.26) would have scored higher than our 6.44. It would still have missed tolerance on two of the three output KPIs.
  • NASA hump: one genuine miss on our side. Reattachment came out 14% late, the over-long separation bubble RANS models typically produce on this case.

HWE-bench is carefully engineered, and these defects were findable only because it makes every number traceable. Reference values deserve the same check.

What a benchmark like this can't see

Even a flawless run through HWE-bench leaves several things unmeasured, and some of them are the reason to build a multi-agent pipeline in the first place.

  • Self-correction is only partly measured. HWE-bench has three recovery cases that hand the agent a deliberately broken cavity to repair. We haven't run them yet, and they're where reviewer agents should show their value.
  • A pipeline's own recoveries don't show at all. On the NACA 4412, our first mesh (426,000 cells) failed our own quality check on 625 skewed trailing-edge faces, and the run switched to a 68,693-cell mesh that passed. The grader only sees the second mesh.
  • Honest refusal can't score. In a product test with no solver installed, the planner cut the run to geometry and meshing, stated that the drag coefficient couldn't be produced, and asked whether to continue. A leaderboard has no way to reward that.

Six checks before an AI-generated simulation number is worth quoting

These apply whether the pipeline is ours, a vendor's or one built in-house.

Put provenance on every reported number. Each value in a generated report should carry the file it came from and the command that re-derives it. This is worth doing even if you never run a benchmark, and it's the check that found every defect in this post.

Run a deterministic reference solve per case with no model in the loop. Without one, a bad score could mean a bad agent or a broken container, and you find out which only after the budget is spent. Check the reference itself against its cited source too, because six of ours were wrong.

Confirm the agent can see every tool the brief names. An installed solver the agent can't find behaves exactly like a missing one, and the agent will substitute something else without the score telling you why. This cost us three cases.

Classify failures on two axes. Record how far the simulation got separately from how clean the paperwork was. The backstep case shows why: a quoting bug and a turbulence modelling failure produce the same number and need completely different fixes.

Build a refusal set. Include cases where the correct behaviour is to decline or to ask, and score whether the system does. No public suite does this, and it's the closest available measure of whether a pipeline will invent a number under pressure.

Record loop efficiency and cost per case. Run the suite with self-correction on and off, and log retries, turns and spend per case. We didn't record per-case cost in this campaign and should have. For scale, SVD Lab's leaderboard reports about $43 per case for a single Opus 4.6 agent on its three-case OpenFOAM subset, and an eighteen-agent pipeline with retries will cost more than that.

Conclusion

  • Make every reported number reproducible. Each value should name the file it came from and the command that re-derives it, and the grader should re-run that command and never trust the report.
  • Keep a model-free reference run for every case. A deterministic reference solve with no model in the loop separates a broken environment from a weak agent before any agent budget is spent. Audit that reference against its cited source too.
  • Label failures on separate axes. Record how far the simulation got separately from how well the result was sourced. Otherwise a sourcing bug and a physics error produce the same score and send you to the wrong fix.
  • Verify the environment matches the brief. Confirm that every tool, solver and file the task names is visible to the agent. An agent that can't find a tool will quietly substitute another, and the score won't say why.
  • Measure what the final file hides. Add cases where the right answer is to refuse or ask, and log retries, self-corrections and cost per case. A leaderboard that reads only the final result misses all of these.

If you're weighing up agentic AI for your simulation workflow and want to compare notes on evaluation, talk to one of our engineers. Our other studies are in the Lab

Building AI into an engineering or product team?

We build AI systems for engineering and enterprise teams. Get in touch and you'll be talking to engineers, not a sales desk.

Get in touch
Samudyata Minasandra

Written by

Samudyata Minasandra

Samudyata is a Software Engineer at Curlscape focused on machine learning and artificial intelligence, with a strong grounding in mathematics. Particularly interested in the mathematical foundations of learning algorithms:
Linear algebra, probability, optimization, and graph-based methods, and in applying them to build reliable, interpretable, and scalable systems.

View all posts →

Frequently Asked Questions

Did a multi-agent pipeline beat a single AI agent on HWE-bench?▼

Not on headline score. On the same model (Claude Opus 4.6), our eighteen-agent pipeline was within 0.07 of the published single-agent run on 10 of 17 comparable cases and ahead on the NACA 4412. It was well behind on three cases where our environment hid the solver named in the brief, and on one case we haven't explained yet.

What does Curlscape do?▼

Curlscape builds AI systems that carry out engineering work rather than describing it. The work falls into three areas. The first is agentic automation for engineering and simulation teams, where agents handle CAD preparation, meshing, and solver setup, along with surrogate models that return a prediction in place of a full solve. The second is infrastructure and data services for organisations where data privacy, cost, or control rules out off-the-shelf AI, including fine-tuned models, air-gapped deployments, and document intelligence. The third is custom production agents wired into existing databases, APIs, CRM, ERP, and document systems.

What does agentic AI mean in an engineering context?▼

It means a closed loop rather than a fixed sequence. The system observes the current state of a case, selects and runs a tool, reads the result, and decides what to do next. The tools are the same software an engineering team already uses, including CAD packages, meshing tools, solvers, scripts, and plotting libraries. What is added is the control layer that decides which tool to run, with what settings, and what to do when the output is unsatisfactory.

Which stages of a simulation workflow can be handled this way?▼

Geometry import and cleanup, simplification, mesh generation and quality assessment, boundary condition and solver setup, execution of the solve, and post-processing into reports and quantities. The deterministic stages are best left scripted. Agentic decision-making is worth introducing at the stages that currently require an engineer to look at an intermediate result and make a judgement.

How is the output verified?▼

Verification is treated as part of the system rather than as an afterthought. Reported quantities are expected to be traceable to the artifacts a run actually produced, and review stages are built into the workflow rather than bolted on at the end. One documented example is a verification agent that flagged a post-processing sign-convention error as physically implausible before it reached the report. This series covers the evaluation methodology in detail from part two onwards.

Does this remove the engineer from the loop?▼

No. Acceptance criteria are set by engineers, and the review of the final result stays with them. What changes is the distribution of their time. Hours currently spent repairing geometry and rebuilding meshes move to the decisions that require engineering judgement, and to mentoring and standardisation, which are usually the first activities to be squeezed out.

Related reading

Latest from the blog

Book a free consultation