Paul Ashcroft (00:01.599)
Hello and welcome to this episode of the Curious Advantage podcast. My name is Paul Ashcroft. I'm here with my co-author, Garrick Jones. Unfortunately, Simon's not with us today, but we are absolutely delighted today to be joined by Dr. Bernhard. Hey Marcus.
Markus Bernhardt (00:20.076)
Hi guys, good to be here.
Paul Ashcroft (00:21.705)
Great to have you with us. Marcus is the founder of Endeavour Intelligence, an independent research and advisory firm focused on AI strategy and organisational transformation. His work starts from a core argument. AI adoption fails when organisations change their tools without redesigning how they operate. Their governance, their decision rights, their operating model. Marcus brings more than a decade of executive leadership and advisory experience to this work. He advises C-suite leaders.
technology vendors globally and publishes original research through the Endeavour report. Marcus, it's a pleasure to have you with us.
Markus Bernhardt (01:00.64)
So really happy to be here.
Paul Ashcroft (01:05.284)
Exactly. Well, the best things come to those who wait, as they say. Well, let's get straight into it, Marcus. Curiosity is at the heart of this podcast. How would you personally define curiosity and what role has it played in shaping your journey from particle physicist to working at the forefront of AI and transformation?
Markus Bernhardt (01:12.312)
Yeah.
Markus Bernhardt (01:25.592)
So for me, curiosity is probably two things at once. it is a personality trait, and I think I've always had it, which is the the itch to know what's really going on underneath. but the part that matters also professionally is I think that it's a discipline where it's a sort of a disciplined not knowing. the willingness to sit with a question long enough to find the real structure.
sort of under underneath it, I think. rather than as we often do, we grab the convenient answer and we move on. So maybe the trait gets you to the question, but the discipline is what kind of keeps you there long enough to find to find a good answer. Not not sure if this is a if this is a good definition. You hear many of these. And for me the the bridge sort of from particle physics, so my my personal background, is the fact that
The thing you can see, the track in the detector, the reading on the instrument, that's never the event itself. It's the real event is inferred from what was disturbed by the particle going through. That's how a detector works, right? It disturbs the measurement device. And the measurement device says, we hear we had something. So you kind of learn rather quickly to somewhat distrust the the reading on the surface. and I think.
That instinct that the surface might not be the actual substance, that's sort of how I also now read organizations and how my my work these days shapes up. and here the I would say the the org chart, then for me, in this example would be the surface reading and the real change is what's happening really underneath in the work before.
Garrick Jones (03:08.412)
Thank you.
Markus Bernhardt (03:14.349)
Before any titles move, before anything in the org chart moves, before someone moves from one team to another or a new new team gets formed. And in a in a hype cycle that we're experiencing, as Garrick mentioned just as we were saying hi to one another this morning, in a hype cycle, that discipline I think is what keeps a a leader honest. You know, curiosity is maybe the thing that asks, is this actually true?
when everyone else in the room is maybe asking how fast can we adopt? so that for me, that for me is critical.
Garrick Jones (03:49.501)
Marcus, you've already said so many things that I'm curious about. not only questions of the truth and the relationship and leaders what they're facing right now. one of the things that Paul, I love the idea that you come from another discipline like particle physics, and there are things that you learn there and things that you
understand from that discipline which you can then bring into other worlds like the organizational world and their insights there. One of the insights that struck me you were talking about traits and discipline and the difference between traits and discipline which are important whether you're implementing AI or whatever you're doing as a leader and then the other thing that really struck me was when you talked about the know the tracking the machine that tells you there's been an event a particle is
gone past is not is not the thing to pay attention to the thing to pay attention to is what it's telling you and that reminds me of that that statement about the map is not the terrain and how it's really important not to confuse the map with the terrain when you're going into new territory or you're trying to understand the context and so on and so forth. find it fascinating particle physics
Do you have other insights from particle physics that you find useful in the work that you do now?
Markus Bernhardt (05:14.44)
gosh t so so many and not just from not just from particle physics, I think from physics generally. When when when you go on the physics journey, what you what you end up doing is looking at a lot of different models. A lot of different thinking models and flowchart models and applying them to different areas and you you become a modeler much more than you become a physicist. I think to be a physicist is to be a modeler.
And to look at how can we how can we model this and how we can we maybe look but below the surface, one or two levels, and find out more about the underlying structures. That's the interesting piece. But I've I've seen these sort of cross-pollination effects in in so many different places. one of my favorite examples goes back to my school days. my my dad taught me to program Turbopascal when I was in second grade.
So I came I came across coding a loop long before I came across the sum or the integral symbol in maths. And when when the sum symbol came up in maths, I thought it was the most fascinating and simple thing ever. It was it was a loop that you couldn't compile or run. You had to think it through them the way you have to think through your code when it didn't work and the compiler just crashed and you go, now I have to check why it crashes and I have to think the damn thing through without running it.
And so and so that is always my example where I I was so far ahead of most others in the class simply because I had come across a concept that was similar but different. And the same thing happened again with the integral symbol, right? The integral symbol or many symbols in maths are just
are just basically code written on a piece of paper rather than in a computer. But you can't run it. You're gonna have to think it through and think through if it would work if someone were able to let it run. But so these examples are all abound. And I would never I when I when I talk about my my observances with what I call the surface wave or what I call the undercurrent, I never I never proclaim to have reinvented the wheel. I'm I'm always borrowing from other disciplines.
Garrick Jones (07:26.47)
There's a way that you talk about the surface wave and you talk about change. And I know that the work you do is fundamentally involved in spotting and assisting with the changes that are coming to organizations. And you talk about previous waves of technological change.
And I also know you talk now about symbolic representation in physics. We could explore so much there, you know, but the importance of modeling and concepts and how understanding concepts gives us insight into things, even as as young people at school. And I know that you work at the intersection of strategy and data and decision making. And the question I really have is, is with all of that.
and your insights from your background and so on, what feels different about the change that organisations are facing now?
Markus Bernhardt (08:26.88)
Staying in the model of waves, the prior waves were about odd adopting a tool, I would say. you bought the thing, you trained the people, and your operating model survived what was happening. this wave is different, I think, because it rewires the operating model. It maybe rewrites. Rewiring is maybe too small.
And here's the thing, it does it relatively quietly, at least in most organizations. So AI shifts tasks long before it shifts titles or the the org chart. that is the that is the really new part. The org reorganization of some of the work has already happened, and we can see that in the data.
but the org chart still looks identical. So the difference is not really speed, it's visibility. And leaders who manage by lagging indicators, by headcount, by job titles, and reporting lines, they're structurally blind currently, maybe to a change that is already underway. And the other thing that data is relatively blunt about is where AI dominates, at least for now.
Dominant pattern is augmentation, not replacement, although that is what we see in the news, right? for every task AI can fully take over, theoretically, that is, right? the the circumstances have to be right. So for every task AI can fully take over, it can augment nearly two others. So the the the the the actual number in the research is 1.9. but the the two to one works really well. And
More than 60% or just over 62% of the work is fundamentally human-led and remains human-led. And so it's it's not it's not people vanishing, although layoffs are often attributed directly to AI. often maybe they're cuts that the organization wanted to make anyway, going through a difficult time, maybe, and and using AI as the excuse is a good way to do it.
Markus Bernhardt (10:42.604)
By the time this shows up in a reorg, the real change, I think would have happened months earlier. And and and that's why and that's why in this in this piece I came up with the calling it the undercurrent.
Paul Ashcroft (10:57.374)
Do think that's what's going on at the moment, Marcus? Do you think organizations are quietly rewiring and then figuring it out? Because I've heard the same, that there is a bit of an excuse to cut the workforce and use AI as the excuse, but do you think the main push is on that kind of rewiring and figuring it out?
Markus Bernhardt (11:18.392)
Two two things there. One, organizations have always had massive cuts from time to time and they are cut they cut deeper than they think they should because they know they can rehire. And we just have to face the simple fact that it's a buyer's market right now. There are many people looking for jobs and it's easier to cut harder than you think and then rehire if you think you've made a mistake. Not quite what we advise when it comes to reducing churn and and what we advise from a talent and HR and and learning perspective, but what is obviously happening.
And then the other thing is, is the rewire happening within organizations? It it really depends on the leaders that are involved and whether they can currently see that it's happening.
And whether they're open with their teams of looking at restructures. We sometimes have big news out there that says, organization so-and-so has now brought the IT and the HR team much closer together, and they think that's how the future works. is that the right way of doing it? We don't know, and they don't know either. And when you talk to them, they're open about that. But what happened is they had they had some senior leaders.
in core positions develop a momentum, develop a shared understanding of where this was heading and at which pace and what they wanted to do. And they decided they were going to put forward a proposal of how to maybe tackle this. And in that organization that landed. So you had the momentum that came from the personalities. They obviously also had enough background knowledge and information to at least sound credible. Probably they are credible. And and so that's where rewiring is happening.
I do see a lot of organizations still that are two years behind and that are not rewiring, that are still working on how do we get our people to prompt better and aren't yet looking at how the technology can be utilized. But that's also normal in in a wave where new technology is coming in, because not everyone is an early adapter. And although the narrative here seems to be around AI in recent years that everyone should be an early adapter.
Markus Bernhardt (13:29.346)
That's never how that's never how reality responds. So I'm not saying that to to make anyone feel bad and say, no our organization is maybe a bit behind. I'm saying this in in in the understanding of no, everyone has to go at their own pace. When I work with a with a Silicon Valley tech company that needs to move at lightning speed because the or the the competition is that's very different to most other sectors. and we have to be aware that every sector has its own
Paul Ashcroft (13:53.661)
Yeah, or highly regulated sectors where it's not so easy to move so fast, right? I mean, in what you're saying, sticking with your wave analogy, I would say there's a wash of tools and new technologies that not just organizations, but individuals within.
within businesses are now exposed to, is available to them. They have co-pilots, they have chatbots, they have agents, they have other generative tools. There is a sea of things to explore and figure out. You mentioned leaders seeing where and how to rewire. Where do you think leaders are misreading intelligence? Do you think they're misreading what intelligence means in an AI world?
Markus Bernhardt (14:46.388)
Intelligence is a really difficult word, and I don't want to get into that definition because that will sidetrack us for the next two hours, probably. The core misread is confusing capability with autonomy, is how I would like to describe it. A tool that drafts beautifully is not the same as a system you can trust to decide. And a lot of what is being sold right now as agentic.
Is really a very fluent chatbot wrapper, coming with a very confident language or a Zapier style automation that has some LLM baked in. But these systems largely have no agency, they make no decisions. It's if-then programming with some LLM around it. And the fact that an LLM rewrites four sentences to make the final email in a Zapir flow.
more personal doesn't make it agentic because there's no agency there. It's very pre-programmed what is happening. And the gap between what something can do and what it is actually allowed to decide is what I call agent washing. I gave a I gave a talk on this at LTUK a month a month ago. So for me, the difference isn't about semantics, it's about non-agency requires an IT governance.
Agency, the being allowed to action something requires a completely new architecture of governance. And the other misread I'm thinking about is judging intelligence intelligence by average accuracy. So the numbers we are all quoting are how well the system does on average or how it does on benchmarks. Every time a new model comes out, we talk about how massively everything has improved on the benchmarks.
but the n the really important number is how the model behaves at the extremes. on the rare high-stakes novel cases, if we want to give something agency, we have to check there and not look at how the average is improving. And systems can look excellent on the average and fail exactly where the cost is high.
Markus Bernhardt (17:03.882)
Or on the case you could not afford to get wrong. And I think that's where a big misread also is. we keep we keep being told that the systems are drastically improving and that they can now self self-correct. There have been many, many such reports, and every time a research piece looks at it, it finds that we haven't really moved that much. The model has gotten better, the average has gotten better, but we haven't moved that much. And the recent Health GPT Mount Sinai study, it was I think it was March.
or late February. Sorry, I haven't got the date with me right now. they they built a model called GPT Health. So they took the newest GPT model and trained it on all of medicine. And then they tested how well it would do. And it did unbelievably well generally. but they also wanted to test how easily it can be tripped up. And the model was tripped up by
even what an eight-year-old would define silly silly context. So as a as a as a rough example, you you give the you give the model the fact that you have breathing issues and you're struggling to breathe and the triage is quite simple. Call the emergency number immediately.
Garrick Jones (18:19.386)
Mm-hmm.
Markus Bernhardt (18:21.76)
If if you add something like and I don't e I don't remember exactly how the study was built, but just roughly, if you add something like a sentence that says, But my uncle doesn't think it's that bad in a in a in a in a in a really crucial amount of cases, the model changed its opinion from you need to phone right now to you should get an appointment within the next three days.
And this is a model that is apparently highly intelligent, can solve mathematical problems that humanity has struggled with for decades, maybe even centuries. all well and good. But an eight-year-old would probably say, Well, we can ignore the uncle for now. Let's let's deal with what we're facing here. And the model the model still falters too often, even in such obvious cases. So
The practical reframe is maybe to stop asking how smart is it? And to really, really, really focus on the boring bit, which is what is it actually allowed to decide and what happens when it's wrong. And b boring but also tedious. That's the that's the non-exciting work, right? We'd like to bring the innovation in and move on. To sit down and really, really build the rules for how we're using it.
and and tackle the question that decides how much oversight you owe the thing. That is the tedious work people don't want to really get into. That's that's the unexciting bit of of innovation, isn't it?
Garrick Jones (19:51.739)
what you seem to be saying is that that's the necessary bit to make it really useful. The boundaries and the rules in place so that it becomes useful for decision making. So it seems to me we're still in a hype cycle then, which is what we mentioned. There's a lot of hype, there's a lot of people trying to market and sell the hell out of these things for obvious reasons and market capitalization is a real objective.
Markus Bernhardt (19:55.406)
Yeah.
Garrick Jones (20:20.808)
You also spoke about the architecture that's changing for organizations and, you know, moving beyond the surface wave of AI to rewiring the operating model underneath. Now, this is fascinating for me. What does that distinction really mean in practice to rewiring the operating model underneath? You said rewiring or even recoding.
Markus Bernhardt (20:46.391)
Yeah.
The so for starters, so the surface wave is all the visible reassuring activity that we normally tend to measure, right? more tools, pilots, prompt training, more vendors in the room with vendor pitches, the dashboards are all green because we have the training in place, we have more tools, it all looks like progress, and nothing has structurally moved. That is the motion without progress piece. And rewiring the operating model means going underneath that and changing.
Probably four things. Who actually has the decision rights? And this can be humans and machines or both. But we have to we have to look at that. Who can decide what and when? What data are we allowed to rely on? Which in IT terms we call data contracts, a technical term that is actually super simple to understand, right? What what data do we rely on when we make that decision?
And who's kept that data clean and up to date. And then the third would be how you oversee it while it's running. We know how we did that when humans did a job.
You know, you had a manager, you you you had spot checks, you had all those good things happening. with systems deciding things potentially, the speed is very different. So a human sitting there and watching it do the work is not how it will work. you know, 500 emails can go out in 10 seconds at 1 a.m. and the human coming in at 9 a.m. and checking what happened is not going to be enough. And and then lastly, the fourth, how the roles are genuinely redrawn around the new way the work flows.
Markus Bernhardt (22:32.738)
Because the human role will change. When once we have more architecture of systems deciding things and moving things and doing things, the the human roles change. And we probably have less managerial and we have more architectural roles, and we probably also have more sort of orchestration roles, which for me means the middle the middle management layer becomes thinner and people are become get closer to the work or closer to architecting how the machine helps with the work.
Or people almost move a step up to the orchestration layer and decide how how we will get the the human machine sort of hybrid system to work. and we kind of already said this. the honest version of this is much smaller and less glamorous than a transformation program. So it is really, when you start, it is pick one workflow, redesign that one process so the work itself.
Is genuinely easier, then prove that it works and then expand from there. Redesign the work first and communicate after the improvement is visible. And the data in the research is also blunt on this. The early gains come from redesigning one process, not from launching a program. And my my example is always when with in an insurance we set out to automate a small fraction of the small claims.
Under 150 US dollars. They set out to automate the first 15% of those, the most easiest. They ended up automating just above 7% after after an interesting journey over three weeks, and letting the II run in parallel to the humans. The humans still did the work. The AI said what it would have done. The data was compared, and we found out that the humans were never as as as
Garrick Jones (24:03.605)
you
Markus Bernhardt (24:27.35)
Consistent as we thought they'd been. Well, surprise, surprise. we found out where the AI did the job better and where the AI fell down because we hadn't written the rule book properly yet. And so these the these are the book and boring architecture pieces, but IT isn't gonna come in and build that for you. You need you need the people who used to sign off the claims to help build the architecture because IT doesn't understand the architecture. And when the system breaks or becomes brittle.
and doesn't work anymore because a change in policy has come in, for example, then again you need the people who used to do the job to help build the architecture. And in in those transformation pieces, I'm I'm yet to come across a team where the team is shrunk.
The ROI is clearly there, but the ROI in none of the teams that I work with has been in people reduction. The ROI has been in it works better, it gives the customers better service. And in this case, the people get told that their claim has been signed off within five minutes of submitting the claim. I mean, that's that's customer service. but the system is still brittle and early on, and we're still architecting how it works, and maybe they might like a look at the next five.
to 7%. Or they might look at the first 5% of a claim that's a little bit bigger. But the the the the workflow in the architecture work is still is still big and critical and crucial. And so yeah, this whole idea that we've handed it over to AI, it's now running beautifully on its own and we've let half the team go. That I'm that I I'm still yet to see that. I I hear these stories.
Garrick Jones (25:55.153)
More hype. More hype. But you make a really powerful point about if you're to tackle a very large system and you want to transform it and you want to do architecture, you say, let's just choose one process and let's get that right. And let's get the benefit of that broadcast so that people can then contribute to the other stuff that needs to be done. I mean, you're famous for your Endeavor report.
Paul Ashcroft (25:55.807)
It's.
Markus Bernhardt (26:24.877)
Yep.
Garrick Jones (26:25.052)
and all the research that's been done behind that. In the last one, when it came together, what genuinely surprised you? Something that shifted your own thinking perhaps from this research.
Markus Bernhardt (26:40.162)
What surprised me is where the friction actually sits. So we all assume the hard part is getting AI to produce the work. And it's not. Production is increasingly cheap, and the bottleneck has moved to deployment, to the infrastructure and the governance around getting that work safely into the flow.
You know, the sandbox piece, the where does it go wrong? Where were the humans better? Where was the AI better? Now we compare the two and we improve the system further. In the research data, you know, one thing that stands out is DevOps becomes the top became the top pressure index. the constraint on growth is no longer writing the code, it's shipping it. The code is coming in at 7, 8x the amount it used to come in because people are utilizing tools.
Often it's called vibe coding, but I give coders still a lot more credit than calling it vibe coding when when when good coders are using the tools. and I would guess that leaders keep reading that friction as a hiring problem when it's really a structural problem. and so the second surprise was that well it was it was a surprise and it was beautiful because
A clean three-layer shape kept reappearing no matter how we slice the data. a layer where the work happens, a layer where deploying it gets blocked, and a layer where the oversight lives. And this held under completely different ways of cutting the data. And when you when something holds like that, it tells you it's a real signal and not an artifact of how you modeled it. And the third one generally changed how I think. as the
Machines do more of the execution. Human judgment doesn't become less relevant. It becomes more premium. a distinct governance and integrator layer in the data emerged with around two-thirds human at its core, as I said earlier, sixty-two percent, just above, is still human-led. And it's often invisible because it hides behind the job titles and the scar the
Markus Bernhardt (28:58.178)
The scarce resource now is contextual judgment, I would say, not raw output. And there's a there's actually a lovely example of this. And unfortunately, Simon couldn't be here, but his team at EY have looked have looked at something that may have looked like a budget constraint, but it was the thought process that forced their strategic clarity. They reframed an LXP implementation.
into a frictionless growth ecosystem. And it worked precisely because the constraint demanded the justification. And sometimes external pressure, budget or otherwise, is the thing that produces the clarity. And you know, Simon Simon lived that one with the team, so he can tell it probably better than I can.
Paul Ashcroft (29:45.971)
What does a frictionless growth system mean, Marcus? You said a frictionless growth. What does that mean?
Markus Bernhardt (29:52.525)
Sorry, say again.
Markus Bernhardt (30:04.119)
It mins
It means that we've we've sometimes we make decisions
Markus Bernhardt (30:15.906)
Because friction gets us into a thought process to solve a problem. And when we come when we come across a solution that really, really well fits the problem, and it's eye-opening, it's like, wha d why have we not always thought about it this way? This is so obvious. This just this is just the way it would work. That's the kind of example where I would say the the the growth was then.
Frictionless because the implementation was so straightforward and everything just nicely fits together. I'm not saying that's how easy implementations can be. I'm not saying they are frictionless. people are gonna people are gonna slam me for quoting it like that. But but I think there's there's there's a beauty when the solution really nicely fits the problem and fits the team and and and things just happen to to fall into place and work to some extent.
Paul Ashcroft (31:12.415)
You've talked about the shift from AI literacy to AI fluency. And when we speak to people, there is a huge range of experience and comfort level with using AI tools. What does AI fluency look like in an organization? And how do you start to get there from the work and the research you've done?
Markus Bernhardt (31:35.754)
Yeah, I I started I started making this distinction probably just over two years ago where I said literacy is foundational and it is knowing what AI is and what it's not. This is, you know, this is when the GPT 3.5s came out and people started saying, I can prompt, I know what AI is. and we need to know what it can do and what it can be trusted with and what it cannot be trusted with. So we need to we need to experiment with the tool and find out what works and what doesn't work.
And so that is the literacy piece. And that is that's why I started making this distinction that this p this is the necessary piece, but it's nowhere near sufficient. And we've been on that journey for several years now. So then fluency is is the applied version of that. being comfortable, utilizing it, and continuously self-improving for oneself with the team.
With new tools that come out, it it it it is more fluent than than just some basic knowledge about how a tool works and what to use it for. And I would in this in this agentic era, I've I'm starting to sort of build a new model within that fluency into into sort of three buckets.
One would be, and a lot of people have been experimenting with this maybe in in in Claude or GPT or Gemini when it comes to when it comes to gems, when it comes to projects, when it comes to so-called agents. One would be what are you building with it? What are you building for yourself with it that helps you in in in on your own laptop, on your own PC when you do the work? The second piece is can you judge what is actually worth keeping and what is actually worth utilizing?
And you will drop some projects and some projects will evolve and be helpful. And the third one is can you can you run and can you operate and oversee the the work that you're doing with these with these tools? That's that's sort of become the new agentic piece because two years ago it was easy. You prompted.
Markus Bernhardt (33:48.31)
You gave it some context, you got answers back, you were the human in the loop, and you decided whether it was useful. Then you went back to an old workflow. More is changing underneath now. And so fluency become fluency is starting to become more to what level are you a builder in your own environment? How do you build your own context? How do you stake token efficient? And how do you experiment around things that might help you in your workflow and share them with the team and decide what's useful and what to bin? And so
Garrick Jones (33:56.422)
Hmm.
Garrick Jones (34:14.844)
Hmm.
Markus Bernhardt (34:16.556)
You know, that's that's for me roughly how I look at it by now. and and but the thing is, as as I'm trying to sort of explain alongside this, this is also in the i in the flow of things, right? the the as as the technology and the tools change, the the definition starts to change. It would it was much easier to differentiate between what I would call literacy and fluency when it was just the prompting world.
Garrick Jones (34:43.46)
Yes, and now you talk about being token efficient, where we were still getting our heads around what tokenization means. Two years ago, now we're looking at making sure that we're the economy of tokenization, it kind of becoming a thing. I want to take us back. You said so many things. I want to take us a little bit back to something you mentioned earlier, which relates to what you were just saying, which was about human judgment and where in your mind does human judgment still clearly outperform
machine intelligence for example and why is protecting that distinction that you made so critical over the next decade do you think?
Markus Bernhardt (35:22.55)
This this is this is almost this is the philosophical one, right? And this is hard to get right. So I'm attempting to give you a good one, not not the correct one here. human judgment outperforms on the novel and on the ambiguous. And also human judgment is still where the accountability lies. We can we can blame the model. Uh-uh. The model made the mistake, but come on, that's not how accountability should work.
Neither in a team nor in an organization as a whole. And that's what that's why governance is such an important piece here. so it's the cases that fall outside of the model training distribution, where context and consequence matter far more than pattern. And that is precisely where machine performance tends to invert, where it is weakest exactly when the stakes are the highest. And I gave you the Mount Sinai example.
Where just the sentence that the uncle said changes the triage of what kind of a medical case you are. Although most eight-year-olds, like I said, would probably say, let's ignore the uncle, let's just make the decision here based on what the facts are. The the system is still thrown by this. And so great examples of where where the human layer is key. And the governance and integrator work stays
overwhelmingly human around these two-thirds that we said, right? And as execution automates, this is the part that becomes more valuable, I think, not less. And oversight becomes the premium skill. And I'm not just talking about the oversight that we used to do in the old prompting world, where the oversight happened in real time, because everything happened at human pace. The oversight now needs to be architect architectured also very cleverly for when things happen at agentic pace.
If I give an agent if I give an agent agency and it is allowed to action something, then the stakes change. If if I if I give a model my credit card.
Markus Bernhardt (37:23.82)
And I say, book me really nice weekend with my wife. It has to be really special because this is a special wedding anniversary, and it books me the best weekend ever and puts 100K on my credit card. then I'm facing a slightly different issue than if I just went in and asked for some good suggestions and I held back the credit card and gave it no agency. And that's the difference. If if I'm in the loop, then I have my human judgment there and then.
We also need the human judgment before we hand the credit card over and say this is really special, spend lots of money. And that's of course a silly example, but th th you know, the second example is, well, what rules would need to be in place for me to hand over a credit card? And if you've worked with an assistant for for quite some time, you'll be happy that they book the flight in the hotel room for you because you have an ongoing you have an ongoing relationship and there's trust involved. trust becomes a different thing when a machine decides. And so
Paul Ashcroft (38:20.883)
We, I think we would agree. We've heard that when you're working with an agent, you must have to think of the agent as a junior teammate or member of staff that, okay, they come with, let's say, maybe your brightest new teammate. They come with lot of intelligence or knowledge, but...
Markus Bernhardt (38:22.274)
That's the difference.
Markus Bernhardt (38:31.01)
Yeah.
Paul Ashcroft (38:41.831)
as you say, they lack a bunch of things that you wouldn't actually trust them with until they've had the training, until they've learned your patterns, your needs and what you need to do. What do you think of that? And the idea of how we are going to interact with agents in the future.
Markus Bernhardt (39:02.412)
I think how we will interact will look very similar to what we're being told right now it will look like, but we're just utterly in unrealistic in the year of the IPOs of these big companies on when it will happen. Right? The Mount Sinai study is a critical example that you would not use an agent for triage. You want the triage nurse to do the work and say, no, here, if someone's having breathing problems, we know what to do.
We're also looking at work again, now that we're on the medical topic, we're we're looking at research around how people make decisions. And it's interesting how our human behavior changes when we know that the second opinion is from a machine. If someone knows that the second opinion on looking at a you know at an X-ray or an MRI, if we know the second opinion is human, doctors behave very differently than when they know that the second opinion is AI.
As they become more confident with the AI, they start relying more on the AI output and are happy to go with it. Whereas doctors who work with other humans that is a second opinion don't don't tend to the research shows they don't tend to change their behavior. They don't become more trusting of that colleague. They see the they see the friction in the system as the whole point of having a second opinion. But when we know it's a machine and it's proven its worth, we start to become a little bit lazy and we say, Well, the machine is probably right, let's move on.
Paul Ashcroft (40:14.228)
Yeah.
Markus Bernhardt (40:25.752)
To the next one. And there's so there's many levels of how we work with these systems. and when we start to trust them and how we build those environments. And we you know, dis despite what the world might look like, we're still researching how humans suddenly change their behaviors in unforeseen ways. And
Paul Ashcroft (40:42.527)
Yeah, do you think so? mean, the analogy that pops in my head is that for many, many years, pilots can land an aeroplane with autopilot and often actually in the fog, they will choose to allow the plane to land itself because actually that's safer. But often they will land the plane, not because the plane can't, but because if they don't do it, then they lose those skills for doing it. But the question I want to ask, and you mentioned this before you said, probably it will give a better outcome.
And is that part of what's happening at the moment where AI, generative AI, of course, by definition is probabilistic. It probably will give you the next outcome, right? But will it give you a deterministic outcome? Can you rely on it to give you the same output each time? And have you in your report studied any organizations that are starting to...
try to get more repeatable, robust processes, which of course are going to be vital in any regulatory, critical, and indeed just business processes.
Markus Bernhardt (41:46.782)
Yes and no, because AI can AI can work very well, especially the good old-fashioned AI that we put into if-then workflows and into decision-making processes that are more like a Zapir flow, where decisions are made with specific data in place, and that that can be extremely reliable.
Through the marketing and through what's been happening the last three years, we start to think that every AI has to be an LLM. And if we try to, if we try to build the same process within an LLM, the architecture is so faulty that we continue to fall down. And the Mount Sinai example, again, is an easy one to come back to. If you build an if then
triage system around someone being breathless, you get the same result every single time. Call the emergency number right now.
And AI can do that. You don't even need AI. That could be a simple if-then with some factors and some decisions. That can just be a decision tree. All those things are still available, but because of all the promise of how intelligent these systems are and all they can do, our minds immediately go to, well, can I build it in OpenAI or on Cloud? And there the problem is that they're the architecture in how these systems are built.
Is non-deterministic and is not reliable. And that's why we see these effects that are not surprising to someone who understands the architecture. It is not that surprising that it hiccups when we tell it my uncle doesn't think it's that bad.
Garrick Jones (43:26.716)
Especially because it doesn't have the context and it needs language context, needs experiential context, needs data context, it needs something that allows it to be specific and also contextually relevant. mean, so they've been experienced.
Markus Bernhardt (43:30.712)
And
Markus Bernhardt (43:38.359)
Yeah.
Markus Bernhardt (43:45.73)
And and just just there. The the the beauty there is that this is exactly the right way to think, but you can't stop there. Because if you stop there, then you would say, well, then we just give it the right context and we build it all up. And and what we're what we're starting to also see is because of how these how these systems work, if you give it all the rules and all the context.
Garrick Jones (44:01.21)
Not enough.
Markus Bernhardt (44:14.462)
Even if those rules in that context is way below what the context window is that these systems can take in.
Markus Bernhardt (44:24.394)
Once you go over a certain amount of rules and contexts, it again starts to become brittle and start to not work anymore. So if I if I you know many models have two million tokens context window. And you think, well, I'm far away from two million tokens if I've given it 150,000 tokens instructions and rules. Well, you try to run an agent on 150,000 tokens instructions and rules.
Garrick Jones (44:32.304)
Hmm.
Markus Bernhardt (44:49.27)
If you give it 150,000, then even if in your first five sentences you've given it five hard rules that it should never, ever, ever, ever fail. Once you get to 150,000 tokens, I bet you that one in ten it'll fail one of your five golden rules that you gave it at the start. So within the architecture, and this is counterintuitive because the system is intelligent, it's AI and it can follow rules. Of course it can, it should. It's a computer program. It's it's completely counterintuitive that
Why when we give it more rules and more details that it then is becomes more more likely at failing again? So what we're really looking at is a minimization problem. If you give it too little, it can't do the job well, and if it gives it too much, it starts falling down. And that has nothing to do whether whether the newest model will give you two million context window or ten million context window or however much they charge you for it. The architecture is wrong in an LLM to make that work.
Garrick Jones (45:43.772)
and does it become a.
Paul Ashcroft (45:44.608)
We, and you saw this, sorry Gary, you saw this, 4.7 and 4.6. 4.7 was worse at creative writing than 4.6. Better at coding analytical tasks, but you know, they're tuned for different things, aren't they? These models.
Markus Bernhardt (45:58.722)
Yeah. And then that's another thing that better isn't automatically better. You would think that if you've built something in the background that works with four point six really well, that when they release four point seven everything becomes a little bit better. And yes, in some use cases, yes, in other co use cases it falls apart. And now now you go back to the team that is trying to run that is trying to run something automatically with this.
Paul Ashcroft (46:04.2)
Mm.
Markus Bernhardt (46:24.44)
Well, the architecture hasn't changed in that company, the architecture hasn't changed in that team. People have just given you a newer model and they've and they and they've turned off the old one that you've been running on. So now you need to run all your tests again to see what has actually changed because it's not just better. So the tedious work is gonna have to be the real work, which is checks and balances and governance and making sure that everything is tickety boo.
Garrick Jones (46:49.99)
that complex dynamic systems tend to much complexity and the system collapses and too many tokens too much.
Markus Bernhardt (46:56.11)
Yeah.
Garrick Jones (46:58.182)
context, not enough context, the system collapses back. Listen, we can carry on talking all day about this. It's so fascinating and your research is so interesting, Marcus. I want to ask a question while we sort of bring this to a close about you, beyond your current work and your obvious fascination and knowledge about AI and what's going on. What are some of the things that you're generally curious about outside of that?
Markus Bernhardt (47:29.198)
Far too much. Far too much. Gosh, how do I pick? no number one, where will it where will it truly change how research works? And med medicine springs to mind first and foremost. How will how will research around medicine, around medication, around treatment, how will that change? And and
If if we make huge progress there, who who will have access to it? we already know that the world is very complicated when it comes to medications and which ones have a patent and which ones are how expensive, and who gets access to the newest and who doesn't. So a lot of a lot of societal changes attached to that. But for me, you know, we we are all humans, and it is it is very vital to us to be.
To be healthy and when something doesn't quite work, to get the right treatment to put us back on track. So both from a research as well as from a societal angle, I think I think the medicine piece would be one of the ones that I'm the most curious about. and and that's why I love I love diving in the that kind of research. And when something like the Mount Sinai study comes out, that's why that's why one of those that I dive dive in with the most enthusiasm. But gosh, beyond beyond that, so many questions we
Paul Ashcroft (48:59.19)
Marcus, it's been a fascinating conversation. Somebody's working at the forefront of this. I've been quite taken by how you've rode back to really the basics of getting some of the basics right in AI. mean, let me try and give a bit of a summary of what we've talked about. And then I'll come to you and maybe ask you if there's one thought you would like to leave our listeners with on this topic about AI and...
putting AI into their work. But we've talked about why this AI wave is different. I've loved our conversation about particle physics and what we can learn from particle physics in the current world.
It's different because it's rewriting the operating model and a lot of leaders are still misreading actually what's going to be required as they're rewiring. You've talked about the surface wave and I know in the Endeavor report you go into more detail about what you mean by the surface wave, but you were giving us some clues around who decides what data they use, who has oversight and how the roles are redrawn in the overarching things that are changing in organizations. I was particularly taken
your view on the role of the human and what I think I took away from this that humans are ever more relevant in the work of working with AI in terms of particularly accountability, governance, how we are perhaps better at dealing with what is novel.
the AI is good at patterns, but not necessarily consequence. And so when it comes down to something that really matters, like should you call the emergency services or delay for a few days, then humans make the right choice. AI doesn't necessarily do so. Congratulations on the Endeavour report. We'll post links to this in the chat. It's a fascinating study of...
Paul Ashcroft (50:50.603)
of AI in global organisations and some of the forefront studies and experiments that are being done. And I think, as you said at the end there, it's about striking that right balance when you're working with these tools to get, essentially in my words, a trusted outcome. You've got the right checks and balances in place to work with AI and agentic AI in particular in a safe and useful way. Garak, anything I've missed, what would you add to that?
Garrick Jones (51:16.508)
Well, nothing you've missed. think that's a brilliant summary, Paul, frankly. The thing that Marcus's work is really making me think about is an incredible book that I read recently by Benjamin Labotard called When We Cease to Understand the World. And it's a small book translated from the Spanish, but he's got about eight chapters about eight different physicists. And each physicist
fundamentally changed.
the notion of physics, know, from Einstein, of course, to Bohr and others. But he also looks at in each chapter, not only the change in the physics that they were using to understand and the mathematics that they were using to understand the world, but what happened to them and the real stories, but what happened to them physically as people in the chaos that surrounded their lives as they transitioned from one system to another system of thought. And the idea that how the way we think
about things actually affects our world. the other thing, mean, I also watched last weekend again, that incredible movie Arrival, which is directed by Denis Villeneuve about the coming of the aliens. But fundamentally, it's about language and how the language that we use
completely transforms our perception of the world, our perception of time and our perception of things. And I know this board is on the philosophical, but there's something about what you've been telling me and what we've been listening to today, which is fascinating about your work, how it gives us insights into what the new world might be and how we need to think about it.
Paul Ashcroft (52:59.079)
And Marcus, so back to you. We've covered a lot. Is there one thought or one thing that you would specifically like to leave our listeners with today?
Markus Bernhardt (53:08.802)
Well for those for those who are who are looking to start the journey, I would say it can all be very overwhelming and the work can be quite tedious because w workflows are quite tedious. So I would say you're you're saving yourself if you s if you start with the small version. So I would say pick one workflow that you own and don't ask where can I add AI, but where has the work already quietly changed? Find the place where the undercurrent has already
Shifted the tasks somewhat inside that workflow, because people are using AI, and redesign maybe that one process. And and that you can do, that you can do with on your own, with your team, with the few people around you. You can you can do that this week. And also if you if you want a structured starting point for those who are slightly more technically interested, I have a free agentic readiness diagnostic.
at endeavorintel.com slash diagnostic, which gives you an honest first look at whether you're running is whether whether what you're running is actually an agent and whether you're ready to govern it. It's a little bit more on the technical side, but it's also written as a bit of a learning journey. And many people have gone back to me and said, I learned so much by going through the questionnaire. It wasn't just a questionnaire for me. so that
Garrick Jones (54:29.02)
Everybody loves a freebie. I'm sure we'll put the link into the notes below so everybody can get to it.
Markus Bernhardt (54:34.254)
That would be my advice.
Paul Ashcroft (54:38.111)
Thank you, Marcus. And where can people go to find out more about you, if they're interested in your work and your research?
Garrick Jones (54:42.897)
Hmm.
Markus Bernhardt (54:44.78)
Yeah, so me and my research, it's it's all on endeavorintel.com. That's where the that's where the Endeavour report sits. That's where I have my briefs that I write every every one, two to three weeks, depending on when I have a thought process that I think is worth writing about and putting out there. and then of course LinkedIn is where I'm the most active and where the ongoing conversation can continues to happen. So please feel free to connect there.
Paul Ashcroft (55:07.039)
Brilliant. And we know you keep that up to date and you're pretty much at the forefront. So anyone who's interested in staying up to date, please go check out Marcus. Well, Marcus, thanks so much for joining us. It's been a fascinating conversation. Really enjoyed it.
Markus Bernhardt (55:19.392)
I enjoyed it as well and thank you very much. And what thank you for the wonderful summary. I was listening to that and thought, my god, that's all we covered. Wow.
Paul Ashcroft (55:24.991)
That's not as good as Simon, I'm sure. But with Simon's listening, I'm doing my best, Simon. You've been listening to a Curious Advantage podcast. This series is about how individuals and organizations...
Garrick Jones (55:26.278)
Yeah.
Great job.
Paul Ashcroft (55:35.549)
Use the power of curiosity to drive success in their lives and businesses, especially in the context of our new digital reality. It brings to life the latest understanding from neuroscience, anthropology, history, business and behaviorism about curiosity makes these useful for everyone. We're curious to hear from you. If you think there's something useful or valuable from this conversation, we encourage you to write a review for the podcast on your preferred channel saying why this is so and what you've learned. Also, we would really appreciate it if you could give us a review and maybe tick that five stars on the five star rating.
it really helps us continue with the podcast and get great guests like Marcus on the show. We always appreciate hearing our listeners thoughts and having a curious conversation. Join us today at hashtag curious advantage. Curious Advantage book is available on Amazon worldwide. your copy today, subscribe and keep exploring curiously. See you next time.
We recommend upgrading to the latest Chrome, Firefox, Safari, or Edge.
Please check your internet connection and refresh the page. You might also try disabling any ad blockers.
You can visit our support center if you're having problems.