We are living in an age of large-scale AI models, and their capabilities feel almost magical. These systems can generate text, code, images, and videos with unprecedented quality, far surpassing anything we've seen before. Their capacity to interpret and generate human-like text has unlocked new frontiers in customer service automation, content creation and decision support.

Despite these impressive advances, AI has limitations that leaders must understand when shaping product roadmaps and distributing resources. This article highlights the strengths and weaknesses of AI to guide its most effective applications. The landscape is constantly evolving however, and we need to stay current with emerging capabilities as this view may change rapidly.
Understanding AI
Before exploring AI applications, let's clarify what we mean by AI and some key related concepts:
Artificial intelligence (AI) encompasses technology that enables computers to simulate human capabilities like learning, understanding, problem-solving, decision-making, creativity, and independent action. This broad category includes:
- Generative AI (Gen AI): Systems that create new content – including text, images, code, and videos – in response to user instructions
- Machine Learning (ML): Models that predict specific outcomes (such as insurance claim costs, property values, or burglary risk) or detect patterns within large datasets
When discussing AI, it’s important to recognise that many people use ‘AI’ to refer specifically to Gen AI rather than the comprehensive definition that encompasses ML. Within the UK General Insurance sector ML models have been industry standard for years and are no longer considered groundbreaking. That said, when working with Compliance teams, remember that the regulatory definition of AI typically encompasses all variants – including ML systems.
What is Gen AI good at?
Below are some examples where Gen AI excels and can be harnessed to drive automation and increased efficiency.
Language Styling – Adjust tone and phrasing, making communication more positive, concise or aligned with specific stylistic preferences. Useful in drafting emails, reports or marketing material and in adapting content for different audiences.
Summarisation – Quickly summarise long documents or information from the internet enabling faster review without extensive reading. This can be combined with triage criteria to enable more effective workflow management. Also useful for turning meeting transcripts into a summary of actions and decisions.
Classification – Categorise documents and images to automate workflows e.g. routing invoices to payment, litigation documents to legal etc. Detect and characterise damage in images.
Answering Questions – Pull specific data from documents, images or admin systems e.g. dates, diagnoses, test results etc. Can be used to extract a predetermined list of items so reducing manual data entry, or in response to user questions, powering chatbots.
Sentiment Analysis – Identify positive, negative, or neutral tone in customer calls, meetings and reviews.
Coding and Documentation - Generate code to fulfil well-defined purposes and create technical documentation like API guides from existing code.
Creative Content – Produce marketing or training materials including social media posts, images and videos.
What is Gen AI not good at?
Despite its strengths, Gen AI has clear limitations. It is important to understand the boundaries, so we do not invest time in product development which attempts to extract reliable results beyond the capabilities of the current generation of AI models.
Accuracy – Gen AI produces ‘hallucinations’ – plausible but incorrect information – because it relies on learned patterns rather than verified knowledge. This can be problematic in contexts where accuracy is critical, such as claims settlement or underwriting decisions when there is no human in the loop.
Confidence Scores – LLMs are overconfident in their accuracy and in most scenarios not capable of providing reliable confidence scores. Leading LLMs can report 90% confidence whilst in reality they are correct in far fewer than 90% of cases.
Numerical Modelling – LLMs are designed for language not mathematics. Because they work probabilistically rather than through precise calculation, they cannot build predictive models to quantify claim settlement amounts or event risks.
Spatial Logic and Complex Context – Gen AI struggles with context-dependent reasoning such as spatial situations. For example, it can miss that a UK front seat passenger cannot hit their right elbow on the door or fail to interpret nuanced liability scenarios in accidents.
Admitting Uncertainty – Gen AI is not inherently capable of admitting ‘I don’t know’ or asking for further information. It tends to generate an answer even with insufficient information, sometimes giving confident but incorrect responses.
Domain Specific Expertise – Gen AI cannot replace specialised knowledge in areas like medical diagnosis, legal interpretation or actuarial science without significant risk.
Other important issues
Understanding available options and their impact on accuracy and cost is essential for informed decision making. Below are more details on key issues to be aware of in planning AI powered product development.
Bespoke Deep Learning Models – These are a subset of AI models which were the only option for language modelling before LLMs became widely available. They are still the gold standard for accuracy in some domain specific areas for tasks like classification and entity recognition. They are also faster and cheaper to run although require some upfront investment of time to train models for each application.
Premium Paid-For LLMs – After bespoke deep learning models the premium LLMs are the most accurate. They offer advantages in terms of speed of development and flexibility as they do not require training (apart from some very specialist cases).
Premium vs Open Source AI – Premium paid-for models are currently more accurate than open source Gen AI. However, the open source models are catching up and we should monitor the performance of these models as they may provide a more economical route to product development in the medium term. In some cases, even today, lower cost or open source models may deliver adequate performance and deserve consideration.
Changing Business Model – It is important to monitor the progress of Gen AI models as they enable businesses to rapidly replicate functionality that has taken many years of development. GenAI models may be able to match the performance of bespoke computer visions models in the not too distant future, disrupting current business models.
Accuracy – Full automation without human oversight demands the highest levels of accuracy and reliable confidence scores. Bespoke deep learning techniques can produce fully calibrated confidence scores, enabling users to automate high-confidence cases while flagging lower confidence ones for human review.
Fine Tuning – LLMs out of the box have good general understanding of language but sometimes struggle with specialist tasks such as medical terminology. In a process called ‘fine-tuning’, additional data can be used to further develop the model to improve accuracy in specialist areas.
Prompt Engineering – This emerging role involves crafting optimal LLM prompts. Despite appearing simple, achieving the best results requires sophisticated techniques. Success depends on rigorous scientific testing, including quantifying accuracy and back testing changes. Prompt changes should be version controlled and ideally stored in a prompt management platform, allowing rapid iteration without requiring application redeployment.
Product Design Principles – Rather than deploying LLMs as monolithic, single-shot solutions, we can achieve better outcomes by creating specialised agents that excel at specific tasks then orchestrating them to accomplish complex objectives. Embracing this modular approach is essential for maximising LLM effectiveness.
Data Capture – Like humans, Gen AI makes mistakes. Build feedback loops into AI applications so users can validate outputs and correct errors. This data provides valuable training material for refining models and improving performance.
Variability – Unlike traditional systems, Gen AI models work probabilistically rather than producing identical outputs every time. Ask the same question twice and you might receive different answers. While certain parameters can reduce this variability, it can't be entirely removed, so proper safeguards are essential to mitigate unintended outcomes. Interestingly, humans share this trait – we're equally inconsistent when answering the same question at different moments.
Managed Cloud Hosting vs Self-Hosting – Managed cloud hosting services (such as AWS or Azure) are much less complex, as the cloud provider manages the computing resources for you, but this convenience typically comes at a higher ongoing cost. Self-hosting offers greater control over data, privacy and infrastructure, which may be desirable in some legal and compliance environments but is typically more complex and requires physical hardware (such as GPUs) as well as technical knowhow.
Gen AI Costs – Gen AI pricing depends on token volume, where tokens are sub-word units of text. Input and output tokens are charged separately, and because tokenisation methods differ across providers, costs vary by model.
Much product development relies heavily on premium LLMs accessed via Amazon Bedrock, which can accumulate significant costs without oversight. Applications should incorporate cost tracking and controls from the outset, and teams should monitor usage to manage spending and inform commercial pricing. As model costs vary substantially, provider selection is a critical financial consideration.
The rise of agentic AI: Moving from conversation to action
The AI landscape is evolving with the emergence of Agentic AI. Unlike traditional AI which mainly responds to queries, Agentic AI executes tasks and makes autonomous decisions. This emerging technology has the potential to revolutionise various industries by automating complex processes and optimising workflows. For insurers this represents a shift from ‘providing answers’ to ‘driving results’ as shown in the examples below.
Agentic AI Applications in Insurance
- Fraud Detection & Investigation: Continuously monitor claims for suspicious patterns, perform fraud checks, cross-check data and identify inconsistencies. Automatically escalate claims to a Special Investigation Unit (SIU) and provide investigators with documentation explaining the referral rationale.
- Automated Claims Processing: Gather claim information, confirm coverage, evaluate a fair settlement, present the offer to the policyholder, and initiate the payment once accepted.
- Dispute Resolution & Negotiation: Analyse the reasonableness of a claim, draft detailed justifications for adjustments, calculate counter-offers, and communicate these directly to the claimant.
- Policy Administration: Handle customer enquiries about mid-term adjustments, generate instant quotes for coverage changes and process policy updates immediately when the customer accepts the revised pricing.
Conclusion
UK insurers already leverage substantial automation powered by AI, particularly machine learning models. Motor and home insurance policies are routinely issued without human intervention, relying instead on customer inputs, automated underwriting logic and model driven pricing. Some insurers have already extended automation to pay claims – Lemonade notably settled a stolen bike claim in just two seconds.
The breakthrough is combining generative AI with existing automation, which enables natural conversations while dramatically expanding what actions the system can perform.
Agentic AI elevates AI from a tool that provides information to an active workforce. By handling repetitive, high-volume process, it drives efficiency while freeing staff for work that requires judgement and expertise.
The distinction going forward will be between companies that use AI to execute tasks versus those that only use it to answer questions. Embedding AI in operational workflows, rather than confining it to document processing and chatbots is where meaningful impact lies.