A New Quality Frontier for Helix
6 min read
“Better is a claim. This is the measurement.”
A new version of Helix is live in DNA today.
Helix is the part of DNA that reads, drafts and answers: the documents you upload when you add a client, the chat panel, the compliance letters, the needs-analysis narrative. It also acts on what you ask, proposing a trust or correcting an asset, and nothing it proposes reaches the file until you approve it. This release raises the quality of what it produces. The largest change is in what it knows.
A new quality frontier
Ask Helix something about insurance or planning and the answer either holds up or it does not. We measure that against answers written by experts. Every question in our benchmark carries the specific points a correct answer has to make, tied to the passage in DNA’s reference library that makes each one true, and we score how many of those points Helix actually made.
The new version reaches 92%, up from 76%.
Put that in terms of a real question. Where a complete answer has four things to say, the previous version reliably said three. The new version says all four.
Where the gain lands, by subject

The gain concentrates where an answer has to assemble several rules at once: underwriting thresholds, product detail, corporate structure, estate mechanics. Tax moves least, and it was already the previous version’s strongest subject.
A score that holds
An average is easy to like and hard to trust. What matters to you is not the score Helix posts on a good day, it is the answer it gives on a Tuesday afternoon with a client waiting.
So we ask every question three times and keep all three. The previous version’s runs landed between 73% and 78%. The new version’s landed between 91% and 92%.
How far the score moves from one run to the next

Its worst day is better than the old version’s best day, and its runs sit inside a single point of each other. That is the difference between a tool you check and a tool you rely on.
It also stopped contradicting the reference, which we count separately from completeness and never average in. The previous version asserted a claim our reference answers call wrong. The new version asserted none.
Where the knowledge comes from
Helix does not answer these questions out of thin air. DNA keeps a reference library of insurance and planning material, and Helix reads from it before it says anything.
We measured what that library is worth by switching it off. Same questions, same independent grader, same version of Helix, with the library taken away.
With the library, Helix reaches 92%. Without it, 71%.
What DNA’s reference library adds

Those twenty-one points are the part of the answer that comes from DNA rather than from the model underneath it. It is the difference between asking a general assistant about a corporately owned policy and asking one that has read the material your regulator expects you to have read.
The subject that does not move is the one that makes the rest believable. Advisory reasoning questions ask Helix to weigh a situation and take a position, and no passage in our library answers them. The library adds nothing there. That is what it should add, and it is how we know the other six subjects are the library working rather than the measurement flattering us.
Reading your client documents
Upload a statement, a policy or a beneficiary form on Add Client and Helix reads it, then fills in the file. The clearest gain here is in the notes it writes.
The previous version wrote too many. Most were not invented. They were real facts told one paragraph at a time, so a single document turned into a long list where a few lines would have done, and the notes worth reading were buried among them. One topic is now one note.
How much of what Helix wrote into a client’s notes was worth keeping

Nine in ten of the notes Helix writes are now notes you would keep. It used to be closer to half.
Everything else Helix pulls out of a document held its accuracy or improved on it: assets, debts, coverage, beneficiaries, client details, and the companies and trusts behind them. Particular mistakes stopped happening. A dependent parent’s chequing account filed as a household asset. An employer recorded as the provider of a group policy. “Estate of insured” entered as a beneficiary. A property-tax notice booked as a debt. Each of those appeared in every run of the previous version, and in none of the new one.
The documents Helix drafts for you
Compliance cover letters, note rewrites and the needs-analysis narrative all start as a Helix draft that you read and edit before anyone signs anything. What changed here is not how much those drafts say. It is how much of what they say is actually in the file.
An unsupported statement is a specific claim the file does not contain. The previous version wrote sentences like “you do not have a mortgage or other debts” where the file said only that one property was mortgage-free, and “both of you aim to retire comfortably at 62” where the file gave two ages and no opinion. They read well. That is exactly the problem. A gap in a letter is visible on the page. An invention is not. The new version writes fewer of them.
Two other things about these letters are checked by a rule rather than by judgement. One is voice: how often a draft reached for a phrase that does not sound like us. The other is the reasons-why letter’s opening, which has a rule about how its coverage paragraph starts, so the letter moves from your client’s situation into the recommendation instead of leading with the product.
Two rules the drafts are held to

Names stay out of the letter
Where a household has more than one member, the reasons-why letter is written to its reader, as “you”. Every version we tested wrote a member’s first name into it anyway, and the instruction not to was already there in front of it.
So it is no longer a matter of instruction. Helix now reads its own draft for the household’s names before that letter goes anywhere. A draft that uses one is sent back to be rewritten. A rewrite that still uses one never becomes a letter.
How we measure this
Three rules, and all three have cost us numbers we would rather have published.
A different model always grades. A model marking its own work is generous in exactly the place that matters. When we moved to an independent grader, one of our own published figures turned out to have understated our reference library by around four times.
Never one run, and three is only a screen. On the same test with nothing changed, one of our scores moved by thirty points across three passes. Any single run would have looked like a real finding. Three passes catch that much. On the test that reads your documents they were not enough either: a change we had measured as an improvement turned out, over more passes, to be hiding a fault that only showed up in some of them. Three runs now tell us whether a change is worth measuring properly. Eight tell us whether it ships.
We never average completeness together with wrongness. An answer that leaves something out and an answer that is confidently wrong are different kinds of harm to an advisor, and one combined score hides the one that matters.
Every benchmark here now runs against Helix on every change we make to it. The reason we can tell you a number moved is that the number existed before we moved it, and that something other than Helix checked the work.
Helix is live in DNA today.
Every figure here is the median of at least three independent runs on DNA’s internal evaluation banks, graded by a separate model or scored automatically against expected results written by hand. Results vary by subject, and the smaller subjects move further on a single answer. These describe how Helix performed on those banks. They are not a guarantee of any particular result for an individual document, question or letter.
See DNA in action
Start building rigorous, client-ready insurance needs analyses, faster.