← All notes

Crikey! Aussies reckon AI adds more work, not less

Have they got a kangaroo loose in the top paddock?

Newsletter #16

Stay around for some news from the land down under

A word from the person behind the laptop

No secret - I love Australia (and my Aussie wife didn’t even have to hold a gun to my head to make me say that).

The people, the nature, the sports (even if I haven't the faintest idea what's happening), and the strong sense of mateship and go-getter attitude. Sure, the meat pies might never compete with Danish pastry, but there's something about the Aussies' no-nonsense, down-to-earth approach to life that I find great.

Anyways, the story of this week is the Australian Senate Committee’s findings on testing out AI to support with government tasks. This is AI without the lens of VC dollars, start-up hype, or hyperscale dreams. This is government officials rolling up their sleeves and hoping LLMs can make their daily grind a bit easier - and this is where it gets interesting.

Because the result shows that it actually does the opposite.

Aussie AI might look different in the future

As a staunch advocate of the idea that “AI change is already here”, it’s refreshing to see someone cutting through the noise and summarise the results of the LLM with a quote like this about their experience:

“It wasn't misleading, but it was bland. It really didn't capture what the submissions were saying, while the human was able to extract nuances and substance. I think that is not a bad summary.”

Short, sweet, and brutally honest. If the ASIC was to use an AI setup like the one in the test it would actually add more work for the employees. Here’s why.

Aussie, Aussie, Aussie: AI, AI, AI

Earlier this year, Amazon put AI to the test for Australia’s corporate regulator, the Securities and Investments Commission (ASIC), using submissions made to a Parliamentary inquiry.

Between January 15 and February 16, 2024, ASIC ran a proof of concept to see how well LLMs could summarise those submissions. What on the surface seems like fairly simple task - but actually requires a lot of attention to detail.

Teaming up with Amazon Web Services Australia (AWS), ASIC set up an proof of concept (PoC) AI project. The PoC wasn’t about making AI a regulatory or operational tool just yet - it was more of a trial run to see what AI could do. The findings were all reported and are publicly available here.

ASIC is Australia’s integrated corporate, markets, financial services and consumer credit regulator

A fair conclusion would be that there’s still work to be done. The LLM-generated summaries lagged behind those created by humans across all criteria that had been determined as part of the project.

It might just be this use case and/or the implementation, but it seems like a pretty methodologically well done analysis, where there was both openness to cut out some boring work by using AI, and the right support to actually implement this in practice.

In my opinion the timeline was perhaps a bit too optimistic measured against the level of perfection needed from a tool like this. Government offices (should) have a low tolerance for errors, and the trial was only running for a month. But to understand how it was done we should look at how it was actually done in real life.

Here’s how ASIC put the LLM to the test

So let’s have a look at how the project unfolded: forget everything about traditional AI benchmarking that you’ve seen plastered all over Twitter for the last few years. For the ASIC

”the objectives of the PoC were to explore and trial Gen AI technologies, to focus on measuring the quality of the generated output rather than performance of the models”

And honestly, that’s a smart move because unfortunately very few people care where your AI ranks on some leaderboard or which model you’re using as long as it gets the job done. It’s all about results, not the tech behind it. Solid start!

The workflow used for the project

Phase 1: Model Selection

The first challenge was picking the right AI model for the job. ASIC tested three different models: Llama2-70B (from Meta), Mistral-7B, and MistralLite. Each model was tasked with summarising public submissions made to the Parliamentary Joint Committee inquiry on the consulting industry. 

Some examples of prompts from the project are: “What does this document say about ASIC?” and “How should conflicts of interest in audit firms be regulated?”

On the surface these are pretty simple, but it turns out, the LLM isn’t great at picking up on those subtle nuances the employees of ASIC would find obvious. ASIC gave each model the same prompts and compared their outputs to human-created summaries.

To ensure fairness, the models were blind-tested - referred to only as Model A, Model B, and Model C. The humans reviewing the summaries didn’t know which AI was behind which output, so there was no bias in the judging. After several rounds of tests, Llama2-70B emerged as the best of the bunch - not because it was flawless, but because it was less flawed than the others. So, ASIC took Llama2-70B forward for the next phase.

Metas Llama2-70B was the model chosen for the PoC

Phase 2: Optimisation

Only a week was set aside for the optimisation phase. The adjustment and fine-tuning of the LLM to see how the adjustment of the prompts to the change of technical settings like "Top-K" (how many potential next words the LLM should consider) and "Temperature" (which controls how random the LLM outputs are) affected the output. What could get the LLM to produce more accurate, coherent summaries.

Like always the challenges that ASIC met during this phase were perhaps not the ones they expected (like it often is with AI PoC’s btw). One big challenge for example was to get the LLM to list page numbers for where ASIC was mentioned in the submissions. Turns out, LLMs don’t naturally track page numbers, so the team had to get creative to solve this problem.

They chose to chunk the documents by page and trained the LLM to associate specific content with metadata - basically tagging each piece with a page reference. 

Temperature explained by Colt Steele

The problem is, this kind of optimisation takes time, and one month isn’t nearly enough to tackle challenges like this. In fact, they ended up having to leave this part out in the end.

Phase 3: Final Assessment

With the AI optimised (or as optimised as it could be within the timeframe), ASIC entered the final testing phase.

The AI-generated summaries were blind-tested against human-written ones - ASIC staff weren’t told which summaries came from an LLM and which were from their colleagues. Each summary was scored on criteria like coherency, consistency, and how well it addressed the specific prompts. In the end, the human summaries scored 81%, while the LLM came in at a disappointing 47%.

So, what went wrong? The LLM struggled with nuance and failed to highlight key points or missed them entirely. It also tended to include irrelevant information, or worse, hallucinate details that didn’t exist in the original submission. One of the major issues was the models inability to capture complex, context-heavy arguments, which left the summaries feeling a bit flat and bland.

Here are some quotes from the ASIC assessors in the report:

Sounds familiar?

That’s because it’s the same challenges almost every organisation is facing with LLMs right now - these things take time to get right.

I’ve got to hand it to the Australian administration for being so open about experimenting with AI, especially in a space where it clearly has potential. But I think a lot of organizations still underestimate the amount of effort it takes to actually produce good results. It might even require your organisation to change an entire workflow.

ASIC’s proof of concept drives that point home. With just a one-month implementation window, you’re bound to hit limitations. But it’s not just about time - it’s about creating the right circumstances. Experimenting means testing, failing, and trying again.

Hopefully, this is just the start, because the potential is there.

In other news…

AI and copyright: again, again, again…

A new study commissioned by Germany’s Copyright Initiative has concluded that training AI models on copyrighted material constitutes copyright infringement under European law. According to the report, infringement occurs at multiple stages, from the collection of copyrighted works to their replication in AI output. This adds fuel to the growing legal debate on ‘fair use’ in AI training, with more experts questioning the legality of current practices.​

‘Duel’ might get a driverless sequel

Aurora Innovation, the Pittsburgh-based autonomous trucking company, is gearing up to expand its operations. The company announced plans to extend its current Fort Worth-to-El Paso route all the way to Phoenix by 2025. This expansion aslo marks a change in the legal landscape. I don’t see solution currently, but perhaps the incremental roll might actually be the solution to make it driverless trucks on the roads legal.

Start-up of the week

This one isn’t strictly an AI startup for legal folks (we’ve seen enough document software for now), but more of a personal recommendation based on a tool I’ve been dabbling with. Playground is an image generation platform that simplifies the process by providing an intuitive framework and enhancing the quality of output by narrowing your input. It’s surprisingly effective. Give it a spin here.

Extra toppings

Talking of image generation: AI is tricking us into doing the boring work…

LLMs for LL.Ms: practical observations on AI, law, and building legal technology. Roughly twice a month.