MHR Labs: Managing AI credibility, creativity and cost
AI is now built into the everyday systems we use to work, make decisions and support people. But as the technology becomes more useful, it also raises new questions about originality, trust and cost.
This article brings together three perspectives from the MHR Labs team to explore what those shifts mean in practice: how AI can narrow the range of ideas we see, how information sources can influence model outputs, and how falling prices are changing the way organisations should think about adoption.
Idea diversity in the AI era, by Chris Judd
Last month, I stumbled across a couple of unrelated articles covering oddly similar topics. They got me thinking about the originality of ideas we get from AI.
The first article by Will Douglas Heaven, published by MIT Tech Review, covers AI’s tendency to return the same answers to certain questions. They give a few examples; simple questions like “give me a random number from one to ten” will almost always return seven, or asking different models for a slogan for a new brand of running shoes will give you near identical responses. The article goes on to talk about a startup trying to train a large language model (LLM) to correct this behaviour, but I’m more interested in the underlying issues it brings up around AI model converging to similar answers.
The convergence goes beyond specific prompts. The second article by Samantha Cole, originally reported by 404 Media, discusses the rise in popularity of ‘Elias Thorne the Lighthouse Keeper’. When asked to write a story, researchers found that AI models gave responses with a strange amount of overlap, often including a character named Elias, or someone working as a lighthouse keeper. Digging into this further, Google Trends results showed a significant increase in searches for ‘Elias Thorne’ over the last year, and a quick search on Amazon shows him to be a prolific author of AI-slop books. This suggests many groups are using simple prompts to make content, with their use of AI tools creating significant overlaps in output.
The research view
There are several possible causes for this behaviour. LLMs generate text by predicting the most likely next word in a sentence. To produce coherent and useful answers, they tend to favour words and phrases with the highest probability of appearing next. Equally, most LLMs are trained on similar sets of data, then further tuned to produce answers that are safe, helpful and easy to understand. When a prompt is broad, these pressures may cause the model to follow familiar patterns rather than something original. This pushes the different systems towards offering very similar responses, as seen in the articles.
Taking this idea further, we can investigate how AI is impacting idea generation. This recent article from MIT Sloan brings together research on the impact of AI use in group discussions. The studies found that individuals working with AI assistance produced more consistent and viable outputs. However, groups working without AI assistance were found to generate ideas with more novelty. Overall, it was found that “AI raised the floor of performance but narrowed the variance in outputs.”
AI can also narrow the initial focus of a discussion, often latching onto the first idea presented without exploring any alternatives. The risk is that once that first direction has been developed into a polished answer, it can feel more settled than it really is, making it harder to change direction or explore alternatives.
What does this mean for an organisation?
Due to AI’s tendency to trend towards common answers, organisations may see fewer truly original ideas over time and a lower chance of uncovering unexpected insights.
The most useful response to this problem is likely to be a set of simple working practices. Start with human ideation, giving people time to explore and flesh out a variety of ideas before introducing AI. Once you have a good set of options, bring in AI assistance to challenge assumptions, identify risks and suggest alternatives.
It is also increasingly important to consult a diverse range of perspectives. Ideas should come from varied roles, experiences and voices that give different ways of framing a problem. It may be worth running human and AI work in parallel, rather than feeding every early thought into the same system. Finally, make sure you review AI-generated suggestions for usefulness and novelty before adopting them. A plausible answer is not necessarily going to be a valuable or original one.
Influencing and corrupting LLMs, by Neil Stenton
Demos, one of the UK’s leading cross-party think tanks, recently released a report: GEO for Geopolitics: What happens when AI and information warfare collide. It looks at how the deliberate manipulation of online information can corrupt and influence responses from many popular online LLMs. The report describes false claims that high-ranking Ukrainian officials were involved in cryptocurrency mining and directly responsible for winter blackouts in Ukraine. The evidence cited by the LLMs, in fact created by Russian propaganda teams, was supported in as many as one in six of the models tested.
While some of these tests were carried out on slightly older and weaker models, rather than the current crop of frontier models, it’s still an important and valid discovery.
Generative engine optimisation (GEO) is to large language models what search engine optimisation (SEO) is to search engines, and it’s a fast-growing sector. GEO looks at the most effective ways of influencing LLM responses from a marketing and advertising perspective. Typically, it will be used by the private sector, but Demos’ report reveals the more nefarious side of how it can be used to seed propaganda.
It appears that making these manipulations is easy, and with a surprisingly fast turnaround. Stuart A. Thompson and Tiffany Hsu’s article in The New York Times explores a similar issue, where updates to Wikipedia and other sites can in some cases appear in an LLM responses within twelve minutes of the change.
When LLMs were first released, their main limitation was the age of their training data. But retrieval augmented generation (RAG) techniques are now well-baked into the systems, allowing quick and up-to-date web searches to make the results are new and relevant. And susceptible to manipulation.
The research view
The geo-political aspects of this issue are disturbing, especially with the nature of some of the false stories. But it also exposes a worrying trend with all future LLM sources.
We’ve all experienced the annoying inflation of traditional search results, especially if you’re a lazy browser. There’s an old joke about how page 2 of the Google search results is the start of the dark web! But having an opinionated chatbot telling you a ‘fact’ that has been influenced and promoted behind the scenes takes on a different level, especially when you don’t have an easy way to click the number 2 option.
As GEO replaces SEO in the marketing toolkit, it’s important for organisations to utilise this technology to make sure they’re still relevant in all flavours of search. But the difference is that this time, the data can be poisoned. This has the potential for an arms race between competitors with subtle defamatory comments on Reddit, or a Wikipedia comment influencing responses from an LLM. Purposefully corrupting documentation could hinder tools and confuse users.
What does this mean for an organisation?
The governance around data corruption and manipulation is new or non-existent, making future predictions uncertain. But this makes it all the more important to ensure employees know when to trust AI output and most importantly when to ignore it.
Demos proposed several recommendations, mostly aimed at governments, which would help restrict which sites are accessed as part of the training materials for LLMs (assuming LLM providers follow suit). But at the end of the day, Wikipedia and Reddit look to be a new source of all truth for these models, which is slightly concerning.
When AI tools are intergrated into your systems, it’s important to use the correct guardrails. For the team at MHR, this means defining governance and controls early, ensuring oversight is maintained, and working within your organisation’s data and security policies.
We should be careful about what information we ask for and trust.
The great AI price compression, by Kevin Slater
The cost of AI is falling, but not for the reason you might think. Two separate forces are compressing prices at once: frontier labs such as OpenAI and Anthropic are racing each other on token efficiency, while a wave of open-weight models, mostly out of China and including DeepSeek, Alibaba’s Qwen, and many others, are pushing the cost of ‘good enough’ intelligence toward zero.
OpenAI’s GPT-5.6 Sol now claims a big efficiency jump on agentic coding work, xAI’s Grok 4.5 is pitched as Opus-class at a fraction of the price, and Meta has launched its first paid proprietary API, undercutting premium pricing by roughly 75%. DeepSeek’s V4 Flash is now one of the cheapest production-grade APIs around, with the Pro Max variant matching frontier coding benchmarks despite being free to download and MIT licensed.
The research view
It’s a mistake to frame this a single price war. A ‘price war’ makes it sound like everyone’s racing to the same finish line, and they’re not.
The proprietary labs are fighting over margins, how to deliver frontier level intelligence cheaper while still funding the next training run. But the open-weight movement isn’t about efficiency at the margins, it’s about removing the cost floor altogether. When you can download an AI model for free and self-host it at benchmark parity with paid options, premium pricing stops being a technical argument and becomes more about trust and support.
The interesting question over the next year will not be which lab is cheapest, but whether the proprietary labs can explain what they’re charging a premium for once ‘smarter’ stops being a reasonable differentiator.
What does this mean for an organisation?
For anyone holding the budget strings for AI spend, this doesn’t mean you should wait for the next price cut. It’s a question of where you direct your investments:
Routine, high-volume workloads (such as drafting, summarisation, internal tooling or first pass classification) can likely move to the open-weight or the cheapest efficient tier for a potential cost saving without taking a meaningful hit to quality.
High-stakes or complex work (such as multi-file reasoning, or anything where a wrong answer is expensive) still justifies paying for frontier models, but that premium should be actively retested and not assumed. Remember: what justified spend six months ago may not today.
Vendor lock-in is becoming a bigger risk than it looks. If open-weight models keep closing the gap, a single AI vendor strategy is a weaker position than a portfolio approach, even if the portfolio costs more in integration efforts up-front.
Procurement and governance functions should expect this to keep moving. Pricing and benchmark comparisons ages quickly in this market, and any internal cost model or vendor comparison is worth revisiting on a quarterly basis, not an annual one.
The organisations that benefit most here won’t be the ones chasing the cheapest cost. They’ll be the ones with a clear view of which tasks need frontier grade reasoning, and which don’t.