← All notes

“All in all, it’s just another brick in the AI wall”

Don’t expect another ChatGPT moment

Newsletter #10

Read further to see the inside of a Concorde

A word from the person behind the laptop

Rebranding linear algebra as “artificial intelligence” may be the most successful marketing campaign of all time - but its also undeniable that AI brings real value.

How do I know?

Firstly, I use it myself all the time - simply because of the fact that it’s useful. Secondly, you currently start seeing results from all branches - legal included - where this new human/computer cooperation is invaluable.

However, this doesn’t necessarily mean we’ll experience another wave like the release of ChatGPT. Making LLMs publicly available changed the entire narrative about AI's usefulness. The road ahead will likely be more stable, but there’s still a lot of work to be done. Today, we’ll look into why this is.

Oh, and we’ll also chat about Nvidia superchips, AGI, and the usual startup stories.

Cheers!

“All in all, it’s just another brick in the AI wall”

I’ve been watching the ongoing discussion about the so-called plateau of improvement for LLMs, especially on Reddit. I decided to examine these arguments for and against.

First I wanna say that I actually don’t think it matters so much at the moment because there’s so much potential in the models as they are right now. Even with their current capabilities, there will be many more use cases for personal, business, and government use. For lawyers, the potential might be even greater than in other sectors - but the pace might be slowing down.

Large language models not learning anymore?

Let's start with this tweet by Ethan Mollick:

“If AI really does plateau at the 60-80th percentile of human ability (no sign it will/won’t), the impacts may be stabilizing. Whatever you’re best at (often what you enjoy most), you’re likely to be better than an AI, but whatever you’re not good at, AI can help fill in the gaps.”

Ethan Mollick is essentially saying that AI still holds significant potential as a tool to enhance human performance, particularly in areas where humans are not experts. This perspective is crucial for understanding the progress we’ve made, even if the performance and improvement of LLMs might be stalling.

So “stalling development” is not the same as we won’t see AI getting integrated into more aspects of our life because there’s still so much potential.

Background on the plateau of improvement

So what does stalling development actually mean?

Gary Marcus, in his article "Deep Learning Is Hitting a Wall" argues that despite deep learning's successes, it's facing significant limitations. The article is around two years old and a lot has happened since, but the general logic remains: current AI systems, particularly deep learning models, struggle with tasks requiring true understanding and reasoning, often making critical errors.

“Indeed, we may already be running into scaling limits in deep learning, perhaps already approaching a point of diminishing returns. In the last several months, research from DeepMind and elsewhere on models even larger than GPT-3 have shown that scaling starts to falter on some measures, such as toxicity, truthfulness, reasoning, and common sense. A 2022 paper from Google concludes that making GPT-3-like models bigger makes them more fluent, but no more trustworthy.”

Even though this prediction didn’t turn out to be true, the conclusion remains relevant. Trustworthiness is still very much a challenge, regardless of fine-tuning or advanced RAG pipelines. Marcus contends that solely scaling up data and models won't overcome these issues.

He advocates for a hybrid approach that integrates symbolic reasoning with deep learning to achieve more reliable and comprehensive AI systems. We won’t have time to dig into that today, but I recommend his article.

The argument against the plateau

Despite the debate, I believe the current capabilities of LLMs are already game-changing. As I mentioned earlier their potential applications in various fields, especially in the legal sector, are vast as is.

While it’s true that there are challenges and limitations, ongoing improvements show that we are far from reaching the end of AI advancements. Let’s dive into the specifics of why this is:

Compute power: GPUs, TPUs, and NPUs are becoming more efficient and powerful, driving the performance of AI models.

Context size: Context size is increasing rapidly

Tokenization: Tokenization techniques are improving, as seen in GPT-4o, making models more effective across various languages.

Multimodal models: Models are becoming multimodal, performing better on tasks that integrate text, image, and other data types.

Data quality and quantity: Data quality is constantly improving, and the availability of multimodal and synthetic data is increasing.

Efficient training: Pre-training and post-training processes are becoming more efficient and automated.

RAG: Retrieval-Augmented Generation (RAG) techniques are improving, enhancing the reliability of AI outputs. More on this later.

Cost reduction: The costs associated with developing and deploying AI are decreasing.

The question for legal teams is whether the AI advancements above outweigh the lack of trustworthiness in responses from a potential legal AI. Or is it worth waiting to see if scaling LLMs continue to improve their performance? Especially because a recent study from Stanford University found that LLMs and RAG tools are unreliable in 17 to 33 % of cases.


I’ll spend the next week looking into the paper above, but knowing that LLMs can hallucinate doesn’t negate their usefulness in my opinion. The current benefits are substantial, and waiting for a perfect model could mean missing out on significant advantages these tools offer right now. Imperfections exist, but their potential to enhance productivity and efficiency is already transformative.

No autopilot just yet

Now let’s talk about commercial airplanes for a second.

The progression from not-flying to flying was obviously significant, but I often hear the argument that progress has been “small” since the development of the jet engine. That’s a misconception. The changes have just been less obvious.

There have been substantial changes to commercial airplanes even in recent times, but most of them aren’t immediately noticeable from the outside. For example, the use of digital controls (fly-by-wire) has replaced mechanical controls in the cockpit, the aircraft body composition has become stronger and lighter, and engines are now quieter and more fuel-efficient, with the Airbus 350 as a well known example.

Comparing the development of LLMs to the aviation industry might sound odd, but it’s quite fitting. Just as aviation experienced a plateau in terms of jet propulsion advancements, AI, specifically LLMs, seems to be undergoing a similar shift in the pace of development. However, it's important to note that just because the fundamental nature of flight hasn’t changed dramatically, it doesn’t mean we've hit a wall.

The Concorde was regarded as the pinnacle of mass aviation, though it wasn’t known for its comfort. Fortunately, the journey from London to New York took only 3.5 hours.

Aviation was already incredibly valuable in its early days. Air travel revolutionised global connectivity, and hopefully LLMs will also transform industries, including law. The key is to recognize the plateau not as a dead end but as a call to innovate differently.

LLMs have seen rapid advancements in recent years, with models becoming larger, more powerful, and capable of more tasks. However, we’re now reaching a point where simply adding more data and parameters might not yield the same leaps in performance.

So, while LLMs might seem to be hitting a wall, it’s likely that we can break through with the right tools and innovation.

In other news…

"Heja Sverige!” Microsoft bets big on Sweden

Someone once told me that Denmark and Sweden are the two countries that historically have been in the most wars with each other. Today it’s mostly banter, but Scandinavian rivalry is beautiful. Unfortunately, it seems like Microsoft put their money on the Swedish horse.

Microsoft is pumping $3.2 billion into expanding its cloud and AI infrastructure in Sweden over the next two years. They’re rolling out 20,000 advanced chips across three data center sites in Sweden, using Nvidia, AMD, and their own in-house AI chips.

At least the Nvidia CEO has a Danish nickname

Nvidia CEO Jensen Huang announced that their next-generation AI chip platform, "Rubin," will be available in 2026. This new platform is set to push the boundaries of AI performance and efficiency. People are excited when Nvidia is involved!

Start-up story of the week

This week, we're looking at Definely, the legal tech startup trying to make document drafting less of a chore - like everyone else (sorry). Founded by a bunch of legal experts and tech enthusiasts, their platform uses AI to streamline the whole process. The idea is to cut down on the mind-numbing parts so lawyers can actually focus on the law. Keep an eye on Definely.

Extra toppings

The logic of LLM scaling is applied to other areas of life:

LLMs for LL.Ms: practical observations on AI, law, and building legal technology. Roughly twice a month.