Video transcript
They’re aimed at being fluent in conversing with you, rather than being concise and explaining when they’re uncertain or they don’t know something. So they will give you answer regardless of whether or not it’s accurate. And no amount of prompting or telling it to verify what it’s giving back to you is actually going to make sure that it does give you something that’s real and true. I’m Tim Street. I write a website called diabettech.com which I’ve been doing for 10 years or so now. I’ve lived with type 1 diabetes for 37 years, and I have spent quite a bit of time investigating how people use large language models, the chatbots that are created by the likes of OpenAI, as ChatGPT, by Google, as Gemini, and by Anthropic, as Claude, for living with type 1 diabetes and assisting in understanding carb counting and insulin dosing and adjusting pump and AID settings. So people are using MDI, they’re using AID. The challenge with all of these is you still need to count carbs, you still need to adjust your doses, and that’s quite a lot of mental mathematics. So what we’re doing here is we’re using those large language models, or trying to use those large language models, to assist in managing all types of therapy. What people are believing with the use of AI in helping them manage their diabetes is that it will be a seamless, easy transition to take a photo of your food, give it to the large language model, and then get a response back that tells you what the carb content is, maybe what the GI is, and how to dose the insulin for that. The reality is a little different from that, because as you may or may not know, large language models are actually a probabilistic guess at what comes next. They’re not a clinically defined method for doing anything. Whilst you might think what you’re getting back is a useful set of information, there’s absolutely no guarantee that it will be. They’re aimed at being fluent in conversing with you. No amount of prompting or telling it to verify what it’s giving back to you is actually going to make sure that it does give you something that’s real and true. And potentially the large language model is hallucinating and not necessarily giving you accurate or valid information. And additionally, with that, it’s not necessarily consistent across doing the same thing multiple times. So in terms of using AI for things like carb counting, and insulin dosing, and acting as a digital endocrinologist, you really don’t want to use the open systems that exist right now. So first of all, remember what it is and know what it isn’t. It is not a clinician, it is not a dietitian. It is a probabilistic algorithm, which means that it is not absolutely certain that what it is telling you is correct. It is guessing at the next word based on a stochastic model. And that means that you cannot rely on the information it produces. When it has any uncertainty, or when it doesn’t know something, it would need to be telling you that. And that’s the key piece. None of the systems that work at the moment are that good at saying, I don’t know. That’s point one. The second point is that these systems are not trained on clinical data. they’re not trained on databases of food, they’re trained on the World Wide Web. That’s a lot of information and a lot of noise. So as a result, when you see a response from a large language model, make sure that you think twice about what you’re looking at. And the third point is: You don’t need to not use them. They are there, they are available. But coming back to the first two points, make sure you consider the response you get back. Remember that a large language model is not a doctor, it is not a clinician. It is an actor playing a role. And on that basis you should probably always get a second opinion, whether it’s your own, or whether it’s from a clinician, or a DSN. I think the direction this needs to take realistically is you need to take a sort of a hybrid approach where you can have a data set that you train a large language model on, using clinical data, using clinical documentation. That means that it looks at that first and foremost. And then when you set the response criteria that you are giving that large language model, it’s more likely to come back with an “I don’t know” response, because you’ve trained it on a specific bounded data set. When you then ask it something and it says, I can’t find it in my data set, you’re able to assess that maybe what it comes back with perhaps isn’t as effective for what you want. And that means that we would need to have diabetes or carb counting specific AI models based off the underlying systems that we talked about earlier like ChatGPT and Gemini rather than going to the sort of open, freely available systems as they exist right now.