Turning Daily Legal Review Into a Governed AI Training Loop

When late-stage AI accuracy concerns threatened the rollout of our legal operations platform, we turned a daily task into a sustainable AI training loop - making legal AI training part of the work, not another project.
Thumbnail

We could generate the tasks, but could we rely on them?

We built a legal operations platform in record time, one that successfully connected it to a search-ordering service and our in-house legal AI. Together, these integrations could order the Title Register, read its contents and generate the tailored tasks needed to set up the file. This was central to our plan to automate setup and reclaim 245 operational hours every week.

Near the end of the build, however, we uncovered discrepancies in the accuracy of the legal AI's findings and generated insights. It had been trained on plenty of examples of more common terms or phrasing, but struggled with rarer scenarios. And if the insights couldn't be relied on, the paralegals would have to manually review and generate tasks anyway, cutting into the time savings significantly.

Context

Our legal platform was built with the intention of removing the need for a dedicated Set Up Team by automating search ordering, title review and task creation.

Problem

Improving those results would require reviewing a much larger number of uncommon and hard to find titles, which could take weeks or months and push the release well beyond its planned date.

Goal

Find a safe route to release on schedule while creating a sustainable way to improve the AI beyond launch.

img
img

hhh

hhh

Focused testing gave us a route to continue with the original release

I reviewed the generated insight accuracy data, spoke with Auction Packs, technology and data teams, and ran focused testing with an experienced legal expert. These short sessions were not intended to retrain the AI. They gave us a clearer picture of its limitations, helped us agree where human review would still be required and informed a practical approach to the upcoming release. Once the immediate route forward was understood, I turned my attention to the longer-term problem.

hhh

hhh

A turning point for a more sustainable solution

Further training required three departments to be available at once, rare titles were difficult to source and legal experts had to step away from revenue generating work. We hit the same problem the original team that trained the AI encountered - not enough time and pressure to move onto the next project, letting the AI training stagnate.

Rather than repeat this slow, stop-start process, I looked for a way to capture their expertise through the title reviews they already completed and incorporate training as a by-product of their day to day.

img

hhh

img
img

hhh

Legal experts could shape the tool through their everyday work

I designed the review directly into the platform, showing paralegals the Title Register alongside the AI findings and generated tasks. They could confirm what was correct, explain errors and add anything the AI had missed. Their feedback would immediately build a clearer picture of performance while creating useful examples for managers to review.

Every review built knowledge for the next

Each AI finding would show a team approval rate, helping paralegals see which findings were regularly confirmed and which needed closer attention. Every confirmation or correction would update that rate for the next time the finding appeared, turning individual reviews into shared knowledge.

Corrections would expose false positives, where the AI returned an inaccurate or irrelevant finding. Paralegals could also add false negatives, where the AI had missed something important entirely. Managers could use both signals to select titles for future training and measure whether later model updates improved the AI in practice.

We designed a way to train AI on more than just Title Registers

Title Registers were only the first use case. The same review, feedback and manager approval process could train AI across searches and other legal documents. Each new use case could remove more repetitive tasks while adding expert-checked knowledge that builds a growing bank of legally reviewed data.

Sitting within an operation already processing more than 2,000 real property sales each month, that data bank could create an advantage that is difficult to copy.

Conclusion

What initially started as a stumbling block developed into a mechanism that supported the wider strategy. While still successfully meeting our original deadline, we created something larger than a release workaround. Great UX made expert feedback a natural part of progressing the case, allowing the data bank to grow through everyday work.