Double AI trouble

Source: Adobe Firefly
24 September 2026
Four weeks from today, the Evident AI Symposium returns to New York. Want to be there or tune in? Register here. We’re also doing a webinar with Nvidia on Oct. 8 about platform architecture and AI’s enterprise impact. Sign up for that one here.
This week: Bank AI is getting messier before it gets simpler. A new look at how lenders are reining in model costs. Plus: is the new model taking the AI world by storm worth the hype?
People mentioned in this edition: Charlie Nunn, Debmalya Biswas, Faraz Shafiq, Sergio Ermotti, Raja Akram, Nayan Rao, Luis Cantero, Aryan Woroniecki, Sean Ringsted, Lauren Sloane, Rakesh Nallavalli and others.
This edition is 1,971 words, a 6-minute read. Check it out online. If you were forwarded the Brief, you can subscribe here.
– Alexandra Mousavizadeh & Annabel Ayles
TOP OF THE NEWS
‘AI IS NOT FOR FREE’
Banks have spent the last year figuring out how AI can make everyday work simpler for employees. The cost of that simplicity, they’re finding, is a whole lot of mess behind the scenes.
You wouldn’t know it from how bankers talk about streamlining work. Lloyds CEO Charlie Nunn this week said agents had cut four steps out of a 12-step fraud detection process. Société Générale said AI would make processes 50% less complex as it chases new AI targets.
But every step AI takes over comes with a load of extra work underneath – the engineering of checks and balances that keep tools on the rails and their answers trustworthy. That’s manageable one use case at a time. As banks race to scale agents though, that hidden work has started to snowball.
Debmalya Biswas, an executive director and distinguished engineer at UBS, explained why in new research this month. Say an agent has to check whether someone is running afoul of bank policy. It might tap an LLM to retrieve the bank’s rulebook, summarize the relevant section and compare the prompt to that summary. Behind each step, Biswas estimates, sits two or three other actions – like checking guardrails or accessing its memory – that add complexity even to straightforward jobs.
Take Lloyds’ four removed steps: Using Biswas’ math, that could mean adding roughly 12 new tasks behind the scenes. Scale that up and it gets unwieldier: DBS’ agentic credit memo-writing tool released last month, for example, automates 70 tasks. Wells Fargo is trying to get agents to handle a 1,100-step mortgage process, Faraz Shafiq, the bank’s chief AI product and solutions officer, said last month (see: “MIA-gents,” The Brief, Aug. 6).
How much of a bottleneck – and cost burden – the background work can turn into is putting a premium on controls that can be reused, rather than reengineered every time a new agent comes along. So far, only a handful of banks have proven they can build those effectively. Of the top 10 firms in the 2026 Evident AI Index for Banks – which we’ll reveal in two weeks (sign up here to get it in your inbox) – seven have shown what their guardrails look like. Only one-quarter of the remaining 40 banks have done the same.
CommBank (last year’s fourth-ranked lender) is one that does. The Australian bank built a “groundedness” check for its chatbot, Ceba, that gauges whether what it’s planning to spit back actually matches the documents and data in the system exactly. That’s now being reused in other tools. UBS, meanwhile, built checks into its architecture that automatically trigger reviews when models take more complex actions. And JPMorganChase built a “Fence” tool that dreams up ways a tool could go awry, tests for them and suggests guardrails the bank could implement based on what it finds.
Done right, systems like these may eventually make agents easier to trust and cheaper to scale. But they don’t make the extra work – or the cost of doing it – disappear altogether.
“AI will be very important for us to take down cost and create more efficiencies. At the same time, we need to recognize, we are shifting a big chunk of cost from one segment to another,” UBS CEO Sergio Ermotti told investors this week. “AI is not for free.”
NEXT WEEK
THE NEXT CHAPTER FOR AI IN LATIN AMERICA

On 1 October, Rohan Ramanath, General Manager, AI Core at Nubank joins us for a virtual roundtable to break down the findings from our latest AI Index for Banks in Latin America and unpack:
- How Nubank, a neobank founded just over a decade ago, topped the ranking
- Why Brazilian banks are dominating the regional AI landscape
- How multinationals like BBVA and Santander are intensifying the fight for local AI talent
MODEL CORNER
JEV YOUR ENGINES
This week, Anthropic and OpenAI both announced new models which the firms claim come with price cuts of 40-50%. New entrant TypeSafe AI thinks it can do better: How about 100 times cheaper
That’s the pitch behind its new model, Jev. While AI labs compete to get their models to spit out beautiful prose more cost effectively, TypeSafe is trying to win over the enterprise with a different kind of model – one that just shuts up. Billed as a revolutionary approach, the model has attracted lots of attention online. In this “Corner,” we look at whether banks should buy into the hype.
Model: Jev
Vendor: TypeSafe AI
Why it’s interesting: Jev isn’t another LLM. It can’t hold a conversation or explain its reasoning. But give it a question and a multiple-choice set of answers to choose from, and it will tell you which one it thinks is right. It’ll also tell you when it’s not sure about the answer – showing you the probability it thinks it’s right – something LLMs notoriously fail at when they hallucinate.
How it works: If a bank gave the model a transcript from a customer call and a list of possible purposes for that call (“lost card” or “suspects fraud” or “failed transaction”), Jev would pick one of the options based on the transcript. Instead of writing out an answer one word at a time the way a chatbot does, Jev returns the answer almost instantaneously and costs next to nothing to do it. It’s more like another type of model – called a classifier – that banks have used for some time now. Those models, though, have to be trained for a particular task. “Jev’s impressive breakthrough is that it generalizes so well,” AI researcher Sebastian Raschka wrote. That means it can be applied to new use cases without costly, time-consuming training.
How good is it: Jev’s makers claim it “can’t hallucinate” since it won’t invent an answer outside of the options it’s given. That’s technically true, but it can still pick the wrong one. In a test of how well the model could classify 3,000 banking customer messages, OpenRouter found that Jev scored 81% to Claude Opus 5’s 84%. But Jev was 13 times faster and cost 95 percent less. Both approaches – the LLM and Jev – weren’t as accurate as a good old-fashioned machine learning model that was tuned to the task.
How banks could use it: For now, Jev looks great for prototyping or small-scale use cases. Before a bank has enough examples in their data to train a traditional model, they could use Jev as a plug-and-play option. “Speed & cost is what got my attention,” Nayan Rao, head of engineering at ANZ, wrote this week. “A lot of teams run basic classification (is this spam, route this to A or B, is this safe) through a full chat model built for conversation. That's a lot of horsepower for a yes/no answer.”
Verdict: Jev, like LLMs, is a black box model, meaning it can’t explain how it made a particular choice. That limits where it could be used, but its cost and accuracy mean banks shouldn’t write models like this – and the competitors that will likely come next – off. “We should be positioning them as an ‘and’ not an ‘or’ in the context of Generative AI,” wrote Christofer Hoff, head of identity, infrastructure, AML and AI security technology at Truist.
NOTABLY QUOTABLE
“I would expect that the pace of AI implementation would be not gradual but exponential. What I mean by that is I think it is going to just multiply every year rather than just flatline, and that's what our hypothesis is.”
– Raja Akram, CFO at Deutsche Bank, at an investor conference, Sept. 23
STAT OF THE WEEK

That’s the upper limit of what some JPMorganChase developers can spend on Claude Code, leaked messages showed this past week. It’s a built-in restriction – that can be lifted if you ask your boss, developers said – that’s part of a new coding platform at the bank called Devspace.
Zoom out: Financial firms have been searching for ways to control spending without kneecapping developers. TIAA, for example, sets token limits based on employees’ roles and workloads. Deutsche Bank does something similar, granting employees access to only tools that are “sufficient for their need,” CFO Raja Akram told investors this week. Even those are fairly blunt-force measures though. More banks are thinking of turning to model routing – the practice of matching a task to the model that fits the job best and most economically – instead. Research from JPMC showcased a tool coders could use to save 43% without sacrificing quality. UBS is using a similar routing tool called vLLM.
Yes, but: Giving the right model the right task isn’t a solved issue. Researchers from IBM showed this summer that a model with pricier per-token costs could sometimes do the jobs cheaper because it could reuse information it was already fed. At the same time, routers sometimes struggle over which model will actually handle a particular job best because a task’s difficulty isn’t always apparent from how it’s described.
BOOK YOUR SPOT
THE 2026 EVIDENT AI SYMPOSIUM

Next month, join the sharpest minds in technology and finance in New York City on October 22 as we tackle what it actually takes to deploy AI across global financial institutions today, and surface the trends shaping what comes next.
Take a look at the speakers announced so far, and check out the full agenda.
TALENT MATTERS
TRANSFORMATIVE HIRES
Santander hired Luis Cantero as head of transformation, where he’ll work with Ignacio Bernal, the bank’s chief AI officer, and Natalia Clavero Hernaiz, global head of AI transformation, to use “new technologies to transform countries, businesses, products and ways of working from a global unit,” he wrote. He was previously the founder and director general of Silbo Money, a fintech startup that let users move money through WhatsApp.
Citi promoted Aryan Woroniecki to be global head of services operating model AI transformation and digitization. He joined the bank in 2023 and, since 2024, had been global head of sanctions operations.
Sean Ringsted is now chief scientist at Chubb, where he’ll oversee the firm’s strategy on AI, data and analytics. A Chubb veteran, Ringsted has held leadership roles across risk, digital transformation and analytics. As chief scientist, he sits on the executive committee.
Lauren Sloane is now head of platform investment performance and acting head of finance for data, digital and AI at Westpac. Sloane has been with the Australian bank since 2021 and served as a director and finance business partner for the customer and corporate services unit of the bank for the last three years.
Rakesh Nallavalli was promoted to executive director and head of advanced capability engineering at JPMorganChase, where he’s “leading the AI transformation across MMC (Mainframe & Mid‑Range Systems),” he wrote. He’s been with the bank since 2013.
Ankita Jain joined Wells Fargo as executive director and principal AI architect. She was previously a lead AI/software engineer at Microsoft and spent time at Bank of America earlier in her career.
Mira Pijselman, head of responsible AI at BMO, left the firm to join consulting firm Ankura. Pijselman led the bank’s first responsible AI function and led an eight-person team, she wrote on LinkedIn. She joined the Canadian lender in March 2025.
IN THE NEWS
CANADIAN WEDDING
TD Bank inked a new partnership with Canadian model-maker Cohere, committing $25 million over three years to “support AI development and adoption.” The deal will send a team of Cohere engineers to the offices of Layer 6 – the bank’s research arm – to work alongside bank engineers on “knowledge management” applications of AI. Together, they’ll “focus on practical value, security and responsible adoption,” the release said. Banks have been leaning into this kind of enterprise transformation partnership more, forthcoming research from Evident will show: The 50 banks we track signed four of these kinds of deals with AI labs and hyperscalers in 2025. The TD-Cohere tie-up makes five so far this year. A common feature of these deals is the promise of co-developed use cases: HSBC’s deal with Google Cloud, for example, will lead to more than 200 new AI use cases over two years, the firms said (see: “In the news,” The Brief, June 25).
Société Générale put a new target on AI-led cost cuts this week: The French bank says it’ll save up to $688 million from AI by 2029. That’ll come from increasing its IT efficiency, improving productivity bank-wide by simplifying processes and using AI assistants to handle more incoming calls, the bank’s presentation said. Roughly $400 million of that has already been “concretely identified…through AI use to date,” a bank spokesperson told Evident. The remaining chunk will come as the bank implements more solutions around the firm. The 2029 goal comes as other banks report realized returns: Santander generated more than $90 million in business value from AI in the first half of the year, and TD Bank delivered more than $140 million through the first three quarters of its fiscal year (see: “In the news,” The Brief, Sept. 3).
Chinese labs don’t need distillation to copy U.S. counterparts, one of China’s top AI labs showed this week. Moonshot AI borrowed a page from the Silicon Valley playbook and rolled out a new version of its Kimi model tuned for use in financial services. The tool pulls data from 10 sources, including S&P Global and the SEC’s EDGAR database, and comes pre-loaded with plugins to Excel. It slashes the time it takes to create financial models from seven days to less than one and cuts in-depth research time from up to 20 days down to two. ICBC, the biggest bank in China by assets, is using the tool, the release said. As U.S. firms call to slow down frontier model development, products targeting specific businesses are becoming a new battleground – Anthropic last week debuted Claude for Financial Advisors, and OpenAI announced Astra for Law. The same battle to win over industries one at a time looks to be coming to China, too.
FUN FOR THE ROAD
EYE SPY
As you read a JPMorganChase document, it may now be getting a read on you, a new patent suggests.
This week, the bank detailed a system that can change the font size, brightness or font type of a document based on your facial expression while you read it. The system accesses your webcam and uses AI to check common facial cues, like whether you’re squinting or your pupils are dilated. If you are, it adjusts how the document is displayed to something a little easier to read.
The moral? Maybe think twice before you pull a face the next time you get sent a memo.

WHAT'S ON
Mon 28 Sept. - Thurs 1 Oct.
Sibos, Miami
Thurs 1 Oct.
The Next Chapter for AI in Latin America, Virtual
Thurs 8 Oct.
From Announcement to Adoption: The Real State of AI in Banking, Virtual
Sat 14 Nov. - Tues 17 Nov.
ICAIF - ACM International Conference on AI in Finance, Milan
- Alexandra Mousavizadeh|Co-founder & CEO|[email protected]
- Annabel Ayles|Co-founder & co-CEO|[email protected]
- Colin Gilbert|VP, Intelligence|[email protected]
- Matthew Kaminski|Senior Advisor|[email protected]
- Kevin McAllister|Senior Editor|[email protected]
- Daniel Shackleford Capel|MD, Banking|[email protected]
- Maryam Akram|Senior Research Manager|[email protected]
- Alex Inch|Data Scientist|[email protected]
- Sam Meeson|AI Research Analyst|[email protected]
- Gabriel Perez Jaen|Research Manager|[email protected]
- Jay Prynne|Head of Design|[email protected]
- Marcus Gurtler|Junior Designer|[email protected]
