Precision Versus Accuracy

Like most people, I’ve been giving a lot of consideration to the state of Generative AI lately, what it’s capabilities are and where it might be going, and it reminded me of the comparison between precision and accuracy... a distinction that I believe is relevant when thinking about what Generative AI, as it stands today, is and is not good for.

Before we go there, however, I want to point out that one of the things that Generative AI does exceptionally well... and we should all be duly impressed by this... literally the more you understand the workings, the more impressive it gets... is pattern recognition. We ourselves, in at least one of our aspects, are pattern matching machines. Recognizing patterns involves a set of exceptionally complex functions that nature has honed in living things out of necessity and the fact that we have (almost accidentally) translated this skill into machines of our own making is nearly miraculous.

Pattern Matching Machines

A lot of what Generative AI does revolves around pattern matching... taking inputs, organizing them into groups and distributing them along lines of similarity, essentially defining the nature of their relationships. Imagine a filter that is designed to gender swap and age faces in photos. By training it on a huge number of photos and letting it figure out for itself what we mean by "gender" and "age" it can create a multi-dimensional map of what generally constitutes a male and female face and old and young faces, as well as other attributes like hair, skin, eyes, teeth, etc... and then transform a photo of a specific face by adjusting along any one or more of those pathways.

Though image and text manipulations are different, the same largely holds true for language. A Large Language Model maps out any number of variables that define a given language and can then transform requests along one or more of those lines, for example, through the use of prompted personas. Feed ChatGPT the basic process and recipe for making beer, let's say, and ask it to repeat it back to you as though you were a first time home brewer... then again as though you were a senior engineer at a pipe fitting company bidding a contract for a major brewery. If it does well, you'll get essentially the same content, just re-worded along the lines that correspond to reader level, experience and technical skill.

Getting the Right Answer

All of this boils down to pattern matching and mapping... and even though we've barely begun to explore the capabilities of Generative AI, it's already remarkably good at both of these. Sometimes shockingly good at it... and already past the point where we're entirely sure how exactly it works.

But the fact is that when most people think of how they're going to use Generative AI, they don't think about how LLMs map meaning along contextual relationships or how Stable Diffusion re-creates images from random noise... they don't care so much how the answer is produced, they just expect an answer... and they expect it to be right. But right is a vague term which can mean different things depending on context.

For the sake of this argument, I’m going to split Generative AI’s output into two primary fields here to make what I think is an important point about "rightness".

precision-2.png

While there is some overlap, In Circle A we have primarily subjective tasks such as the creation of music, images and video where "surprising" results and variation are highly tolerated, or even desired... whereas in Circle B we have more objective tasks such as solving a math problem, producing a working code solution or diagnosing a disease, where we would imagine there is a much smaller set of correct results. You might roughly think of them as "art" versus "science" distributed along a spectrum with some overlap in the middle.

Generative AI makes extensive use of pattern matching in both areas to great effect and that’s got people investing billions of dollars in further AI development. But I think it's important to distinguish between the two different types of outcomes represented in the two circles and state that, while I believe we should be very excited about Generative AI’s emergent capabilities in Circle A, we should be equally concerned and cautious about its prospective mastery of Circle B.

Getting To the Point

This where we come to the topic of precision versus accuracy. This chart below should tell the story of the difference between these two concepts pretty clearly. Precision is an indication of the ability to effectively reproduce an outcome whereas accuracy is the indication of whether or not that outcome meets the stated requirements.

precision-3.png

In the lower left quadrant here, we see both low precision and low accuracy. Assuming the bull’s eye is our objective, our “shots” are neither hitting the target nor grouped closely together. The “answers” we’re getting for the question of “hit the bull’s eye” are neither precise nor accurate.

In the lower right quadrant, we’ve achieved precision, but are still lacking accuracy. Our answer is reproducible... we hit the same area every time... but it’s not a bull’s eye. Some people might respond to this situation by simply moving the target (look, we were right all along!) but that’s not how reality works. Precision alone isn’t enough.

In the upper left quadrant, we see a high degree of accuracy but low precision. All of our hits are distributed around the bull’s eye, but none of them are direct and there’s a considerable spread. You might say this is close enough for horseshoes and atom bombs, but if the stakes were high enough, we’d choose the upper right quadrant.

Here we see both precision and accuracy. Our answers are tightly grouped and they are dead center in the target. Now while this might give us a memorable Ted Lasso barbecue sauce moment in a game of darts, the irony is that perfect precision and accuracy are not always ideal. Let’s go back to our Generative AI circles and see why.

It's (Not) All Subjective

As I mentioned above, Circle A represents a set of outputs that are more subjective. When it comes to producing music, images and video, there is no single right answer. If I ask Midjourney to make me a picture of a cat in a space ship, there might possibly be an infinite set of varieties of that image, many of which I would gladly accept as being "right"... of meeting the requirements of my request.

precision-4.webp

They should all have something resembling a cat and a space ship in them which gives us the accuracy component, but if Midjourney gave us back a hundred images with perfect precision, they’d all be nearly identical which would be awful. For the types of tasks that we deem to be creative, a certain amount of distribution of outcome is desirable... and we’re willing to forgive the AI if all of the cats are different shapes and colors and one of them looks like it’s made of folded paper. That feels creative to us and that’s one of the reasons why image creation with Generative AI has become so popular.

The problem here is that sometimes we need more precision in our output. While it may be fine to go through four or five iterations of a text prompts to find something that strikes our fancy for a single image, if we're trying to produce a 3 minute video with consistent characters, backgrounds and color palette, we're going to want much more control over the precision of these elements... especially when changes cost time and money for the system to re-render your entire video clip and the output changes each time you prompt it. (I'm willing to wager that "AI generated commercial" you saw recently underwent many of hours of human production work to make it look like it wasn't touched by a human at all.)

Accuracy Matters

Now let’s move into Circle B where precision is much more important. Let’s say you have a column of numbers to add up and instead of just wanting the sum (where precision and accuracy must be 100% but which could be done in any spreadsheet), you actually want to see some sample work processes. So you ask ChatGPT to show you three different ways to “do the math” to get the result. Accuracy is going to be vital here... each process should get the same result... but precision is less important. You actually want the methods to be different and the more-so, the better.

The same could be said to be true of generating code. Precision wise, there could be a dozen different ways to code a particular function such as “send an email when the form is submitted” but they should all be accurate, in that they all deliver the same expected result... your customer gets an email when they submit your form.

So far we’ve talked about examples where some degree of spread in precision and accuracy are acceptable. But what about medical diagnosis? Here’s where we start to realize the limitations of Generative AI.

While we’ve seen instances where Generative AI was able to match patterns found in sample x-ray data, for example, well enough to identify issues in new x-rays being read, when it comes to a medical outcome, in order for a Generative AI to be considered better than a radiologist we’d expect both a remarkably high level of precision and accuracy. It’s not good enough for a Generative AI to precisely identify a healthy pancreas as cancer cells every single time nor is it useful for it to identify actual cancer cells accurately if it can only do so 50% of the time.

The Gorilla In The Room

For a more real-world Generative AI fail of precision versus accuracy, remember back in 2015 when Google’s image recognition tools started identified particularly dark skinned persons of color as being gorillas... to the point where they had to remove “gorilla” from the system’s recognition set, resulting in “gorilla blindness” where it failed to recognize gorillas, despite correctly identifying other primates like baboons and orangutans.

While this was a tremendous embarrassment for Google, not to mention producing considerable outrage among the affected public at the time, it wasn’t particularly life-threatening. But in an environment where industries have decided to invest billions in “AI” driven decision-making tools, we could start seeing mistakes that impact customer and patient outcomes which rise to that level or worse.

Confounding Outliers

And just to make this interesting, let’s throw in two more variables that further complicate the problem... Generative AI can’t be properly trained on things that, by their nature, rarely happen. Typically when we find outliers in our data, we throw them out... but there are outlier cases which do occur. If the training data for an “AI” doesn’t include enough instances of an outlier condition, the system might not be able to recognize it and it may just disappear into the noise.

The other variable is that, due to the fact that Generative AI works largely by predicting the most likely matching result for each step in its calculations, outcomes are going to tend toward the median. This means that as we run out of organic data and Generative AI continues to be trained on “synthetic” data and the output of other Generative AI systems, more and more of the outcomes are going to move closer to the mean and the overall set of possible “answers” that Generative AI gives us will, over time, be reduced. A unique twist on an Orwellian dystopia... one where it's machine generated entertainment and convenience that constrain thought rather than an overreaching government. Huxley would be impressed.

The "I" In LLM Stands For Intelligence

Enthusiastic futurists believe in the best possible case; one in which our current approach of throwing more computational power and larger data sets at Generative AI will result in it spontaneously evolving into an actual machine intelligence. That's the road we're currently on and the justification for all the hype and the enormous valuations of the current crop of "AI" focused businesses. I must admit I am very skeptical of all that, nevertheless we appear to be running at full speed toward handing our jobs and the creation of our culture over to Generative AI and I, for one, am not sure I want to participate in a culture that is even more artificial than the one we've already made.

When it comes to future pathways for Generative AI, though, I'm not sure which I believe is worse... really complex machines that fake intelligence so well we forget what real intelligence is, or spontaneously emergent Artificial General Intelligence that develops so quickly it evolves beyond us before we even realize it exists. In either case, we can only hope to be well kept once our world no longer belongs to us and hope the machines idea of a "good time" for its human pets turns out more Huxley than Orwell. To be honest, though, I'd rather we went with Herbert and banned all "thinking machines" before it comes to that point.