# AI Agents, Product Discovery, and the End of Human Judgment

**Podcast:** Stories Connecting Dots with Markus Andrezak
**Published:** 2026-07-15

## Transcript

Stories Connecting Dots, a business podcast to discover a world of possibilities through stories told.
Yeah, so hello and good morning, Elze.
Hello, audience.
Happy to see you again, Elze.
This is more or less a spontaneous recording we did because I posted on something, you posted about something very similar and we exchanged a couple of messages and I wanted to introduce the topic a little bit.
Where it originated from is both of us were looking into these new fancy tools that developers get, which is about long-running agents, totally autonomous.
And, you know, Boris Cerny basically saying how I'm working now is I'm not even prompting anymore.
I'm building loops that build the prompts for the agents.
Super cool.
So Elsa and me, we're sitting in front of our...
terminals, cloud codes, whatever we're using, and like, hmm, I'd like to do something like this.
And we can't, by the soul of my mother, we can't find a thing in our jobs which would merit by doing what Boris does.
So I wrote a little thing which was about like, yeah, that's basically because software engineers were working for checks for correctness and all that stuff.
linters, syntax checks, type checking, even compilers, which are basically guaranteed or even architectural tests for the code, which basically can build a loop which is checking against your other loop, which is building a thing.
So you have a loop that is building a thing and you have another loop checking for correctness.
And it stops when everything is correct in a simplified sense.
So we as product managers, managers, whatever we're doing in what we think is more abstract stuff or not that tangible stuff, we're trying to do the same thing.
And of course, we're failing because we cannot define the second loop, which is checking for correctness.
So there is no halting criteria for the first loop that creates a PRD, a strategy document, finds the next feature to build, whatever, because we have not worked on a definition of correctness for a thing we're working on.
which software engineers did for six years or something.
They also built compilers, you know, all that stuff, which kind of about correctness and these things and determinism and whatever.
And that leaves us with a question, which is really interesting, which is it leads to us as product managers coming to the conclusion, not specifically Elsa or me, and that's the thing we want to debate right now, that we say that that's a huge part of marketing of product management right now, that the last line of defense against these small little beasts of machines that are out there is judgment and taste.
And that's the last thing I want to say before I leave the stage for you, Elsa.
And I think this is dangerous territory for a job to be in.
And I just want to give a very small example.
So I write a PRD and I'm...
I did my work.
I worked on it maybe even for a month.
I did my research and it's totally resting on user research and what I discuss with my colleagues.
So I write down the PRD and say like, this should be the solution direction.
These are the tests we're doing, whatever.
Along comes my boss, the PMO.
And what he says is like, oh, I don't like the PRD.
I think I could do a better one because I think I have better judgment and taste than you.
Yeah, maybe like.
But I don't think so.
I think my taste in judgment is better.
So because we're now in the uncharted territory of none of this is deterministic, none of this has to do with logical reasoning or anything, because it's something out there which is more close to art than science.
You know, there's a world out there where all of this is justified when you're writing books, doing art.
The question is then again, like how much of what we're doing should be art and how much in the real world is art?
are 90% of product managers out there paid to do art, basically.
Judgment, taste.
And maybe the last sentence, there's this one epitome of art for me, and I cite that a lot.
It's Gerhard Richter, the most expensive modern artist of our times, and he's standing in his studio.
They're making a documentary about him.
And he's standing like this little old man standing on his ladder and doing these, I don't know, 12 types of gray with insane stuff in his huge atelier.
And the journalist is looking at him and says, like, is it done?
And he's looking to the journalist like, are you insane?
It's not done.
Of course not.
And then 10 minutes later, he's climbing from his small little ladder and he's looking at the journalist and says, like, Now it's done.
And look at the pennies.
Basically looks the same as 10 minutes before, but now he's convinced it's done.
And that's art.
Gerhard Richter says it's done.
There's no argument about it.
You can't even argue about it.
Who would the journalist be to say, hey, Gerhard, why is it done now?
It's just insane.
And the question is, do we want to be in that territory?
And I don't actually have an answer.
It's just like this whole area of...
What are we actually doing?
What is the job?
What does it mean for the technical tools we're doing?
All that stuff is super interesting to me.
And it's opening up right now because there are tools we cannot use, which are totally attractive to use.
So, sorry for the long intro, but maybe that triggers something in you.
That triggers so much in me, I don't even know what to do.
So I'm going to start with maybe taste, right?
I like taste, judgment.
product sense.
I don't know if they're the same, but I feel like they're in the same kind of camp.
To me, it's funny, right?
We were in this whole, in 2015, we all had to be data-driven.
Everything had to be data.
And then 2020 was like, okay, data-informed, right?
Like, not fully data-driven because data can lead us down the right path.
So data-informed, right?
And now it's like, oh, AI can do that.
So it's taste and it's judgment.
But when I think about what are these things, To me, besides meaningless, right, I think it comes from having seen a lot, right, having put the reps in.
And ideally, if you're building for a consistent target audience over many years, you have done a lot of testing and you have collected a lot of evidence from that target audience and understand how they behave over time.
And therefore, your judgment and taste is not this fluffy, arty thing, right, but it actually relies on.
pattern matching or a robust mental model of how these people actually tick and how they might behave and what works in the market and what doesn't.
Right.
So in that sense, I do not think product sense is maybe the better word because taste and judgment are really fluffy, but data or following a structured product discovery approach are not at odds.
Right.
Because I feel like if you spend years.
following a structured approach and therefore collecting evidence and testing assumptions and learning and building this mental model of how a certain market works and how a certain type of user works, that gives you a shortcut to the right answer, maybe later.
So in that sense, I can live with it.
What I find very interesting is this idea of saying, okay, this is human, right?
Human judgment, human taste, and an AI cannot do it, which kind of implies that we would be good at it as people.
Whereas in my experience, we are awful at it, right?
Like I have been in tech for 15 years and I've been training product teams for the last four years as a solopreneur.
Firstly, most companies don't do discovery at all, right?
As you have already said.
So there I'm thinking, Please, if we have an autonomous discovery agent, please.
Right.
Because then even if it's subpar, if it's not perfect, it's better than nothing.
Right.
Which is usually the baseline.
And secondly, it's not like whenever a human is involved, the output is excellent.
Right.
Like that, that guarantee.
So I'm almost like, maybe it's fine if it's an okay-ish autonomous agent doing it versus a human, because yeah, lots of thoughts.
Maybe.
just a short stake in the ground.
Like when you say like, we're not actually that good at it.
Like Kohavi spent a career on it and the numbers he comes up with.
And that's a famous story and study he made about Microsoft.
Like nobody needs to like Microsoft, but of course they're not the worst company on the planet.
objectively speaking.
So it seems they make a lot of money, they know how to do business and all that stuff.
So the study he made, and it's really famous, is like when they are testing features, the numbers are that roughly a third of the features show positive outcome in value creation.
A third of the features show no measurable economic outcome.
A third of the features show negative impact on value creation.
That seems to be rolling the dice to me.
And that is Microsoft.
So that's really interesting.
But these are also features that have actually been tested, right?
Yeah, that's the point.
That's at the last stage gate.
That's not what I start thinking about.
That's like I filter it and filter and talk to my colleagues and then I build them and then I test them.
And out of all the candidates, I think, which are test worthy.
Yeah.
Still, it's rolling the dice.
Like, what the hell?
If that's human judgment, I don't know.
Right?
If we accept that as a starting point, which we as product people should know about, right?
And we also know that product discovery has never been the focus in most product organizations, right?
It's been like 80% delivery focus.
Before it was Scrum, and now it is bi-coding, right?
But this whole delivery theater now has a different code, but it's always been the thing.
Thank God if we can have something that helps us do that and maybe autonomously does it.
It's not like we're replacing people because people haven't been doing it.
And on the other hand, regarding human exceptionalism and so on, and the discussion on how good is that slop stuff and so on that you, Marcus, created, the problem with the numbers that Cohavi comes up with is the barrier of improving human judgment or being better than human judgment is not that high.
Yeah.
And that's something that I always see, like when people are like, you know, we're doing all the stupid experiments, like, you know, we haven't automatically created PRD from user research, blah, blah, blah.
And then people are like, yeah, you created it in five minutes.
It can't be good.
And it needs to be sloppy.
Yeah.
And then the comparison is against like the perfect human with the perfect PRD, which never exists.
And I'm a consultant.
Like I am going into companies who come to me because they complain because product is not working well.
And now all of a sudden that AI is here, it's like, oh, what you create with AI is slop.
So basically it means like what we create now would be good.
It's not.
People complain because product departments are not as successful as they should be.
But all of a sudden, now that AI is the enemy, we're like, oh, that's slop.
That AI stuff is slop.
But what I created until here without being successful is great stuff.
Like, what the hell is the comparison?
Like, what are you doing?
It's really strange effects coming in.
And I think a lot of it is just psychology.
Yeah.
I mean, I also agree with the AI slop argument.
I think there's space for both arguments in my head.
I see like, but we can, that's a whole different topic.
And I wanted to go deeper because...
I started this conversation with you when you posted this thing about, okay, my engineers are doing all these cool, like they're all night long, their machines are running and they wake up and the app is ready.
And I'm here invoking my skill and then checking the output.
And I saw this picture, I wrote this article in February about, can I manage a swarm of agents?
And I have this picture of Peter Steinberger with like 13 screens.
And I'm here like with two.
And I'm going insane with like context switching.
And it's like, I'm doing it wrong.
I'm doing it wrong.
So I got really into that.
And I didn't want this to stay this abstract thing of like, could it be?
But I said, let me take a slice of product management work that I think is important.
And let me try.
to see how far I can get.
Having little experience building autonomous agents, but I have built an agent before and I've built LLM-based workflows.
So I had like a little bit of knowledge and I know the product works.
So I was like, let's see.
And I took this slice of, which I think is generally considered very human work, which I've also studied quite a lot, is I have a bunch of transcripts and I have like a bunch of qualitative data.
And out of these transcripts, I'm going to plug...
opportunities first on a transcript basis.
Then I'm going to cluster these opportunities and then I'm going to size them and I'm going to pick a number one opportunity to focus on.
And I find this sitting at the heart of product discovery.
I've also loved Teresa Torres's book Continuous Discovery Habits.
I've also been in her opportunity mapping course.
So I've gone quite deep on that and I've been doing it a lot with teams.
And I know that it's also one thing that humans find really hard.
And before I started doing an agent, I thought, let me first create these as skills so that I can babysit them, so that I can really manually invoke them and check the output.
And when I think it's good, I can think about, can I convert this somehow into an agent, right?
Because now the theory, the rules are there.
So I didn't get further than the skills at this point.
So I have these skills that I'm kind of happy with.
Right.
I'm like, they are useful.
And people are writing me and posting about them and people are writing me like they are not bad.
They're actually quite helpful.
But I say this is a second set of synthetic eyes.
Right.
It's not autonomous.
It's supposed to help you, but you should do the work.
And then I was thinking about, OK, now let's say I turn this into specialized agents and there's an orchestrator handing off.
What do I need to put in place for this to go in this loop?
Right.
In this kind of loop where it has.
a clear definition of done.
I think that's ingredient number one.
I need to know what is a good number one opportunity, right?
When is it done?
And then when I have the definition of done, from there cascades, what are the checks that I can put in place throughout this cycle that are cheap to run and that can see if I'm progressing towards the definition of done?
And maybe there can also be one at the end that is not cheap to run, right?
Or that can take a lot of time.
And I did this mental exercise.
I didn't build this, but I just did this mental exercise of what could that look like.
And the first thing I stumbled over was my definition of dumb here, which is what is a good number one opportunity?
And my first thought was, well, an opportunity was good if I picked it and built a solution around it.
It was a solution that got traction.
That's a big problem, right?
And I thought first the problem is time because...
it could be six months until it gets traction.
And then I thought, actually, that's not a problem because I can have an agent that just goes to sleep and just checks in in six months, right?
So I think time is not the issue, but I think here attribution is because between something in the right solution and having a product that has traction in the market sits so much, right?
It's like, okay, the opportunity was right.
Maybe the solution I came up with was wrong.
Maybe the execution was wrong.
Maybe the timing was wrong.
Maybe the channel was wrong.
So there's too many factors.
So it's like, okay, it can't be that.
And then I said, okay, maybe my definition of done, it has to be a defensible number one opportunity based on rules that I codify, right?
If you follow these rules, it's probably good, which is not really what I care about as a product manager.
But then I thought, isn't that what software did as well?
And there, maybe I hand over to you, like, is that because software also stops.
Software also doesn't say, okay, the loop is finished when the product is loved by customers, right?
It just goes like, okay, it functionally works and it looks fine, basically.
Yeah, so maybe first on the experiment, I did the same experiment.
So my trigger was that for my product management with Cloud Code course, which is actually not about Cloud Code that much, but like how do I work as a product manager with all these tool skills, whatever, like what is a new style of working with it?
And there was one.
client in the course and say like, yeah, you know, I can't live with all that slop.
Like I need to have really exact results because all my stakeholders on the business side and on the technical side are really, really painstakingly looking for correctness and everything.
And like we're in a small niche of cooling cars.
So they build small series of cars which guarantee a cooling chain for food.
So lots of regulations, anything.
So I built a thing which is running for one and a half, two hours and goes into all angles of the business case and the technicalities.
And I showed it to him.
He's like, that's insane.
That's perfect.
The outcome is so, so good.
How did you do it?
And there was a lot of stuff like the technique I use, like how would I build this as a human workshop?
The normal social problem is that you come to a conclusion too early so that you have early commitment to, you know, half-baked results.
And then I introduce a lot of friction into it and let the person in the workshop create more friction and more friction and get random friction and all that stuff.
So that's why it takes long and it's all these tokens up.
But the results are painstakingly deep and everything.
Getting data that is just deterministic.
Yeah, it's like breaking the workshop basically down into a million steps and each step just gets a small piece of the context.
So it's basically the mirroring what context engineering does for that kind of work and all that.
So I'm proud.
Super cool stuff.
Now I say like, oh, I build it as skills which are calling each other and which then create that workshop kind of setting, whatever.
And it's like, oh, now that it works.
I just build it into an agent.
Like there's these great things.
Agent, like, sure, let it be autonomous.
How cool is it then?
And that sounds so cool.
As you say, like, I'll have it run overnight.
I come back, it does the thing.
How cool is it?
So I build it as an agent.
And then I started with a different business problem.
I come back the next morning.
It works as an agent.
Perfect.
Same result.
The problem is I cannot use the result.
And why?
Because I cannot judge if the outcome actually is that great or not.
Because like when I built it at skills, I sat in front of the machine and, you know, I did my things.
I had my coffee.
I don't know.
Maybe I chatted with you or whatever on the side.
But all the time the process was running in front of my eyes and the LLM, the cloud code was chatting to me like phase number one is done.
And the CTO came up with that argument and blah, blah, that reasoning.
the business, the growth guy needs to come up with new numbers and like, oh, that's the friction I was looking for.
And then I'm sitting there looking at the agents, how they're chatting with you.
And I know how the outcome is created over time because I witnessed the process of the thing.
That's what it does qualitatively when I set it up at Skivs.
When I set it up as an agent in the background, I get a result and maybe the result is perfect, but I don't have the experience of how it's built.
Now my problem is in the next meeting, I would be the person who has to justify that result, but I wasn't witnessing how the result was built.
And here comes what you said before.
Why is it happening that way?
Why can code be checked and this stuff cannot be checked?
First of all, what I said before, like we have...
25, 30 criteria which the code can be checked against.
And then there might be another 10% which are like, do I feel like it?
Is it the patterns we use?
You know, stuff like that.
And the worse the engineering department is set up, the less of these 25 to 30 checks they are doing.
So, which is also...
raising havoc with them now with AI.
Like when they push all that stuff to PR, but they don't have the infrastructure to test all of that immediately.
They either push it to production, it's not done.
But why can it be checked?
Because code is a statement of now.
I can check if it does what it's supposed to do now.
For all the stuff you and I are discussing with our colleagues, We cannot check if it does the thing now.
What we can check is the PID that you wrote.
Is it exact with regard to the statements we know about now?
Because does it serve what customers said, demanded, were ordering?
Yeah.
But what we cannot say about it, and that I think is a huge gap, is the forward-looking time arrow.
Will this thing do what it's supposed to do?
Just because people said they want it, will they pay for it?
That's not conclusive ever.
And I think that is one of the fundamental things which are different about the things that you and I are working on versus what engineers are working on.
Engineers working on a thing that is finished now.
It's finite in this moment and it's only to be checked against the past.
Does it do what the specification says?
Yes or no?
But isn't that totally artificial though?
Because actually software developers and product managers, we all have the same goal, which is build a software product that people love to use, right?
That's it.
We have the same goal.
It's just that software at some point said, hmm, I can't really check that part.
Let me make an artificial cut and stop it at the point of it works according to spec.
So we as product managers could do the same thing.
And that's the move that I made psychologically is saying, I'm going to make a cut and I'm going to say, So basically, like we say in software development, if I build a feature that responds to what a couple of people have said, and that looks the way that the designer specced it out because we agreed that that's good, then the feature is done.
And we're not going to care if people use it.
So we can do the exact same move in product management, in a sense, right?
Because we can say, like, okay, we just cut it at, if it comes to the number one opportunity.
I have this thing of like, okay, we have a couple of checks, right?
We say, do these opportunities, like when the AI in one step extracts opportunities, he has to have direct quotes.
And then we can have a check that says, are these direct quotes real, right?
Deterministic check.
And then at some point you could do this clustering and then you can have maybe an LLM as a judge that checks, does the cluster make sense?
And then you can have, in the end, I have a scoring, which I codified where I say, look.
only look at things that have an importance range, opportunities that have importance range of four up, like on one to five, it has to have at least four, and then look at prevalence.
So there's like a scoring.
And then you can say, okay, that you can game that, right?
But you pick the best rule that you think is a predictor, and then you kind of stop there.
Because that artificial move, software development is making it as well, right?
So I don't, I think it's just a choice.
It is.
Actually, there's a concept and a word for it.
Russell Eckhoff branded it in the 70s, 80s.
I don't know.
It's called Idealized Design.
So the story is the following behind it.
So insanely good book.
Everybody needs to read the book Idealized Design by Russell Eckhoff.
He basically tells the story about the Bell Labs.
So for everybody who doesn't know, the Bell Labs were this, you know, the phone company.
And what happened at Bell Labs was that it was totally unsuccessful.
The boss of Bell Labs looks past, I think it was in the late 60s, he looks past for the last 30 years and says, we exist here, heavily funded.
We didn't do a great invention in the last 30 years.
We did all the relevant inventions in telecommunication were done around the turn of the last century.
Since then, no innovation, although we have millions and millions of budget.
What's going wrong?
So he went back to his managers and what he said is, In the morning meeting, gentlemen, the telephone system of the United States broke down last night.
Bad news.
And the managers were looking around at each other and saying like, no, it's not true.
I called my wife.
It works.
And he's like, yes, metaphorically speaking.
What I'm saying is we need to create conditions where we are now innovative again.
And how we create these conditions is we imagine that the phone system broke down and we imagine a completely new phone system.
which works today.
There comes your point of, isn't all of that just artificial?
And what you said is like, I do not allow you to think about the future anymore.
I want you to take into consideration what we have, the constraints of today, and just work on stuff that works today under today's constraints.
The opposite thing would be the Segway, which was insanely great, whatever.
It just didn't work under the conditions that were given at the time.
It required infrastructure.
And now the outcome was everything we know today about telecommunication was invented in those years after he made that invention with his manager's caller ID.
You know who's calling you.
Invented then.
The Tasten telephone.
What is this?
Keyphone.
invented at that time, answering machines invented at the time, like even in the cloud, so to speak, like that's working with the phone company.
All of that stuff was invented because of that invention where you said like, let's not talk about it.
And that's something I do with all of my clients.
In workshops, I do not allow people to talk about the future because like that is just rhetoric.
And that's basically the interesting part is.
That's taste and judgment.
And what it's ending up with is like only rhetorics.
Basically, the most outspoken person.
normally a man like me, like, would go to an Elze like you and say, like, ha, ha, ha, Elze, you don't have a clue about the future.
I have a better, I have the better crystal ball than you, Elze, because I'm so much more informed.
And then basically, in the same room, no one has data about the future.
That's an interesting thing.
Like, we do not have data about the future.
Any projection is just fantasy.
And the rest is rhetorics, like.
I have a better clue and intuition about the future than you, Elsie.
It's just insane.
But that is at the core of it.
Like how much of the jobs we are doing is that human exceptionalism individually or as a species where we say like, oh yeah, we actually know about what's going on in the future.
Well, we don't, you know, it's just the made up stuff.
Really the job that I've been doing as a product manager is I always...
And this is the thing that I can totally automate, I think.
It's basically, and this is why I set my definition of that a bit earlier, is I follow preconceived rules that I think make sense, which have to do with, for me, like, how do I find the number one opportunity?
How do I ideate multiple solution ideas?
How do I bring it down to a top three?
How do I match assumptions?
How do I test assumptions?
These are all processes where I follow rules and have been following them forever.
And what it leads me to...
is something that I think is defensible.
And then I make a decision.
And this is the stopping point, right?
It's defensible.
And then I make a decision and then we go into future territory and then we see if it works.
And then I get feedback and like, did it work?
And that flows back into my rules.
Like it didn't work because of this.
Maybe this rule was off.
Actually, maybe importance level should be weighted higher, right?
I changed the rule and I kept doing it.
I think this is how my process has been working.
And I think if you do it many times, you can call that judgments because you can skip a few steps and you can skip ahead and you don't have to follow every rule, maybe.
I don't know.
But in the end, you can set your definition as done in the future, which makes no sense, but you can set it here.
And then it depends on how good these rules are.
But these rules get better as you see more future state and you close that feedback loop to what works.
So basically, that's what I've been thinking about.
And that I could codify, I think.
But that brings me to my next step.
And I wanted to ask you about this because you understand engineering much better than I do.
And I have been comparing, again, this flow of teasing out opportunities and then clustering them and then picking a number one opportunity.
I like this one because it's so human and it's so hard, right?
It's so much interpretation of what did they mean.
So I find that a really fun chunk to work on.
It should be a defensible number one pick.
And then I spelled out what that is.
And then I have these verifier checks.
And then I compared it to the verifiers for a definition of done for software.
And I saw that for agentic engineering, it's mostly deterministic checks in there, in these loops.
With a very thin layer of semantic at the end, which is like, are the code patterns good?
Is the plan solid, right?
There's like a little layer of semantic in there, but most of it is quite deterministic.
And then when I look at what I'm doing, it's mostly semantic and a little bit deterministic, right?
So it's like, what do they mean?
Is this the same opportunity as what the other person said?
Or is it distinct?
Or is it a sibling, right?
Super.
And the funny thing is I've done this opportunity mapping course and I've been teaching it a lot.
where I sit in a room with other people who are deeply interested in this and are doing it at work and are trained to get good at it.
And we all pulled out a different experience map, slightly different opportunities.
We clustered them differently.
So these are all people who are really into it and we're not doing the same thing.
And I've been wondering, firstly, is that a feature or a bug?
Because maybe it's actually a feature.
So one thought.
And then secondly...
How can I, should I even, try to move my semantic checks to deterministic?
Is that what I should be doing?
And if so, how?
Or is that actually not what I should be doing at all?
And I just wanted to get your thoughts on that.
Yeah, I think that is so interesting.
I think how we perceive those things might change the more we use LLMs.
in that work so i think a couple of years ago we would have said like semantics are not that deterministic and what you would see with engineers is like when i was working at ebay and all the engineers were fighting about stuff over lunch and coffees and whatever most of the fights were about semantics not that much about like yes like do you use camel case how do you call your variables like what's the naming scheme like you could fight about this but most of the fights are about the The stuff that feels less deterministic, the semantics.
Like what I'm doing now with LLMs is that I start to realize that most of the semantics in a company, in product management strategy, is not that much art as we think, but it's much more deterministic when they think, like, as you say, if it's not anchored in a quote, it doesn't exist.
And I think, like...
epistemologically like how do we get to knowledge i i think product management especially um is trying to get vague and we confuse two things one is correctness the other thing is completeness and like when we think about innovation is like i think the correctness although we didn't accept it is easy to prove like can you anchor what you say, what you want in something you heard somewhere?
And then we call it storytelling and any of these creative techniques and design thinking or anything.
But when somebody tries to conquer you and win over you, basically they come up with something which they think you did not see.
So it's about completeness.
What they say is like, I actually heard, they don't say it like this, it's just a screaming argument or anything, but I think what's behind it is like, I heard a different signal somewhere.
I picked a signal up somewhere which would change everything drastically so you don't have a clue because you didn't ever pick up that signal.
I think that's the thing.
And maybe, but then the question would be like, and I think like if you think about decision theory and stuff like this, how should things be decided in a company, in a social group?
It's like different phases.
First of all, this information intake.
We try to get a...
as complete picture as possible of a thing from different sites that why diverse groups are a good thing like because they pick up different signals of the world of the thing that is and collect them and try to paint a picture from all sides like a 360 camera you know doing all these things and then you have a better picture and then of course you need to time box that because it could never end and that's not how a company works you can't think about stuff for five years or anything like that's what you need to time box Then the next step is actually trying to complete that picture.
And then it's trying to get a lay of that land, the picture that you drew.
And then if that's the case, this is what we're doing.
And I think why we try to make this a more mystic and archy thing is because actually humans are bad at all these techniques.
The thing we have is like how we do sense making of user interviews is like.
We go through the interviews again and again.
We basically retell the story of the interview.
And what it does with a group of people is we now have collective storytelling.
We remember the same arches of the story again and again.
If I show the same interviews to an LLM, the LLM is done with that in three minutes.
But we need hours and hours of repeating and repeating it because one person is picking up that angle, the other person is picking up that angle.
Based on life experience, one person is judging this as more important than the other person.
And it's all cool, it's all fine, but we need to realize that we're just bad at this.
And we don't have like the logic systems in our mind to basically see semantics as a correct thing.
We're like, no, it should be this, it should be that because none of us has the complete picture and is valuing different things.
And it's fine, it's human, it's cool, but it's not better intrinsically.
And if you look at companies who work with this constructive, like, you know, the famous example, Pixar, like they create a wonder.
What they're basically doing is like, Every year at Christmas, they have this movie out there, which is chart-breaking.
Every year.
And the opinions go all kinds of directions.
Like if Pete Docter is not involved, the main storyteller, it's a bad movie.
But every year, it's just chart-breaking.
And every movie they make takes five years in development.
So it's a really, really predictive process they have.
Every movie takes friggin' five years.
Every movie is chart breaking, whatever.
And they created a process on doing that.
And how they do it is basically a feedback process which eliminates the human flaw, which is like, this is very subjective.
They try to create an objective process out of subjective storytelling, which is repeat, rinse, repeat.
Every week, the new storyline is presented to a couple of people which can give all kinds of feedback so that you pick up more and more signals.
The rule is that only the team working on the movie is judging the feedback they're getting.
They pick up what is important to them or not, but they have exposure to all of these signals.
They pick the signals which help them right now.
Pete Docter cannot tell them anything if he's not part of the team and stuff like that.
And that over five years creates a process which is totally deterministic, which is like, we always create a great movie.
It's not random at all.
Like by creating all this serendipity and randomness all of the time, we create a process which is totally stable, reliable, that creates an insane movie over five years because we created that culture around it, which says like, I'm nothing, the movie is everything, the feedback, the signals, the very diverse signals are everything.
So innovation then, The soft part is just about completeness.
And there is actually a technique for product managers out there.
I never studied it that deeply, but the original jobs to be done framework is just that.
It's not this, you know, what is the job to be done?
It's a guy sitting there doing interviews for months, picking up what are the jobs that people need to do in front of the screen or whatever the job is.
And then he's ending up with a wall with thousands of post-its, which is like, These are all the jobs we could think about.
That's how he's picking.
So it's basically a scientific method of mapping the world that is relevant to the job and then clustering that and then saying like, out of all the signals we got, as I said, prevalence, this cluster over there is mentioned 995 times.
This other one is 830 times.
This has more impact than that one.
So we pick this one first.
And then we will do a good job.
And that's insane.
I think when we introduce art to the job, it's because we are not doing the grind.
So the jobs to be done guy, Bob Moster, he's doing the grind.
And I think that's how we need to see things.
It's coming back to what Kent Beck said.
90% of...
Our skills are losing their value.
10% are increasing 10,000 times, whatever.
The 90% is what can be replaced by grind because these machines can grind like crazy.
And that's just the painstaking, number crunching, deterministic stuff that they are so good at, the pattern recognition and the counting and all that stuff.
So you have two thoughts on that, actually.
Because one thing, like when you were talking about the Pixar movie, I was thinking about this.
Pre-note takers and pre-AI, right?
I sat in a lot of interviews.
I really like this example because we're always like, this is superhuman, right?
This whole opportunity extraction and sitting in humans and sitting in interviews.
And I sat in these interviews and I had the luck with product trios, right?
That that was facilitated.
So we had three people or two people in an interview.
And at the end, we compared notes.
And they were always different.
It was always like somebody wrote, I hate it when this happens.
And somebody else wrote, I don't like it when this happens.
And somebody else wrote, this happened, right?
So there was like varying degrees of, did the guy hate it?
Did he just tell us it happens?
Did the guy say he didn't like it?
These are very, we each took notes and they were like slightly different as in like, what did the person actually say, right?
And there I think.
AI is just amazing because it just says, this is what the guy said.
And there's no, we're like, yeah, but you have to read between the lines.
Maybe we should stop trying to read between the lines, right?
Let's just listen to the actual words that were said.
Maybe it's actually better than us trying to see like, oh, but he looked away.
And that's like, maybe let's get the voodoo out and just listen to the exact words that were said that might be better.
So then I agree with you that AI is not just capable, but better.
than us at analyzing, like just objectively better at analyzing, especially vast amounts of qualitative and quantitative data over time.
We suck at this.
We can only hold the last interview in our head and poorly at that.
Even we're misremembering the last thing we heard 10 seconds ago.
So we're just awful at it.
And AI is amazing at it.
But when you said about the completeness thing, I think we have sucked at this completely.
And I think we're humans.
where I see maybe my role maybe going, right, is more on the data collection side.
Because I've always collected, as you say, a very incomplete set of data because I had just a couple of interviews with a couple of people.
But now I might have way more time to, like, for example, when you talked about Jobs of Vida and I was thinking about ODI, which I liked a lot, which has this same idea of complex, as in it walks people through every single step on their job map.
And then...
You know, I did these ODI job mapping sessions with people for like four hours where I bought them breakfast and I had them walk me through every single step so that I didn't just see the steps where they had a ping, but I saw all the steps.
Right.
And I think this is one way to identify hidden needs, latent needs.
So I think this is one thing that is totally underestimated by product management when we do interviews where we ask them about the stories, what happened, people tell them about the stuff that is on fire for them.
And then you have these problems that are solved a thousand times over.
But there's not much, not many product managers know how to collect hidden needs or latent needs, right?
This is like an art form in itself that not many of us know how to do.
I think the ways, there are ways to do it like ethnography and like these ODI mapping sessions.
These are amazing techniques.
Most product people are not using them.
So I'm like, what if my job shifts to collecting data, not analyzing data?
Because I know I stuff at it.
But collecting it because I think it's harder for an AI.
Yes, you have AI interview bots, but I feel like for now, I can maybe get more out of people because it's human to human, right?
It's going to be better conversation.
I can meet them in person.
I have a physical body.
What can I do with my physical body, right?
So that is helpful.
Well, I can sit in somebody's office and I can watch them.
So I feel like maybe my job will shift away from the data analysis towards collecting more complete.
data.
I'm dying right now because all the things you're saying are so spot on.
Like, if you think about ethnography, like, and that's my number one complaint, you know, all these, and I don't want to bitch around about.
product managers like that's not the point like we live in environments which don't allow us to do a better job so what i'm saying now is not against these people so everybody could write a really perfect prd if we had time nobody gets time so it's sloppy what it's so it seems so anyways when we talk about interviews what everybody's going in there with scripted interviews and then they're like reading what people said not what's behind you know what's the need behind it we don't even teach the vocabulary enough of a want versus a need and all that stuff.
So the job that really learns how to do this is the job of the etognover.
And what he does is basically, and that's the funny irony, we might talk about it as an art, but what he does is like, he tries to be as objective as possible.
He tries to not inject himself into the observations.
He tries to be just a note taker.
This is what people are doing.
He doesn't go in there with it.
It might mean this or that.
He's just note taking.
He's just observing, writing down observations.
If you watch a person with that education, you are dying in envy over that person of how much that person can reduce itself from the conversation and just ask totally neutral, open questions in that situation.
Not injecting himself as a person at all.
Just like, and then what does it mean?
And then what would you do next?
And what's happening then?
And what did you do then actually?
And what happened to you after that?
And he comes back and I have two examples of major insights in my professional life which can exemplify how bad we are at this.
And the techniques we need to employ to come to an insight which to other people might be totally obvious.
So the first one was at Mobile.de.
So we were, CEO comes to us and says like, we need to have that B2B app for our dealer.
So what Mobile does is like, you can look for cars, you can then find a car at a dealer and then you can go to the dealer and buy it.
Like there's no transaction on the website.
It's just like, you find a dealer, you go there, fine.
And we realized that dealers also, of course, are buying cars.
And there's kind of lots of cars traffic in Germany because VW Golf might be cheaper around, I don't know, where I come from, Rheinland-Pfalz, than it is in Berlin right now.
So people are trading over the whole country.
So that might indicate that there's a cool cause for a B2B app because dealers are dealing with dealers.
Cool stuff.
Also, the salespeople all come along and say like, oh, yeah, we should build that stuff.
Heavy, deep research into that.
So what we did is we rented to campus and went to Frankfurt where in one road there's around 50 car dealers and we made eight iterations per day of a B2B app with, you know, drafting it, changing it between interviews, whatever.
Then we didn't find out enough.
Then we invited dealers to an on-site meeting and we did innovation games with them.
Like, you know, I just learned innovations games, so I needed to employ it there.
And basically, we were able to ask two questions on the evening of day two.
One question was like, okay, we still understand you all want to have that B2B app.
Yeah, yeah, yeah, cool.
Yeah, yeah, yeah.
And then the second question was, who of you would actually publish their cars on that app?
Because now what you're doing is, dealers to dealer is smaller prices.
You don't sell to a dealer at the price as you send to an end customer, an end consumer.
You don't want to sell from there.
So you build that B2B to B marketplace, basically, where everybody has demand, but nobody wants to supply.
Yeah.
And we're like, okay, so everybody of you wants to have that marketplace, but nobody wants to send their supply into like, how is that supposed to work?
And that is so frigging obvious.
Like the thing is because we were so blind in the business we are doing, we couldn't see it.
We didn't have the completeness.
So we spent days in Frankfurt, two days more in Kleinmachno basically to find that out.
And that was then the insight that led to all the need behind that is all they want is cheap supply.
Where do they actually get cheap supply?
That's from private sellers.
So if you insert your car, as an ad and it's cheap enough they give you a call and say like yeah you know you wanted five thousand for it but you know this specific model with that build year it's going to rust in one year.
I know what's going to happen to the clutch in half a year because you already have 150,000K on it.
So what about I give you 4,000 and you're like, okay, fine, D, I get the cash right away.
It was just inserted for 15 minutes, got it.
So that's the day basically we decided that we need to make private ads for free because that is actually bringing.
Anyways, long story, long winding story with one insight.
Don't build that thing.
do something completely.
All they want is supply.
I bet you with agents running for two hours, you find that out in two hours.
You might not believe it.
You might, you know, because human flaws and everything, but that thing would be on the table very, very quickly with agents would say like, oh, you know, that's a totally unbalanced two-sided market.
Nobody will go.
But we needed ages and ages to find it out.
Same thing with my hammer, like my dark here at my hammer.
Like that's the same thing, kind of two-sided market and transparency on the craftsman market.
And consumers are asking for, could you widen my whole place?
Like it might be a student community and they ask for rooms to be widened.
It never really took off.
It was kind of an okay business, but it was always limited.
We could realize that only bad craftsmen are on the platform.
The good ones are not.
And we were totally blindsided on why that happens until completeness.
We made interviews, stuff like that.
The insight actually came out of not design thinking, not all the interviews, but me spending evenings drinking with these craftsmen.
And I'm not drinking, but anyways.
I was even driving with them on their cars and like, how do you acquisition?
And they're like, acquisition, what do you mean?
I'm like, how do you get customers?
Oh, okay.
Okay, wrong vocabulary and stuff like that.
Like all this traditional stuff.
But the only moment I learned was, and it sounds amazing, but it's a flaw, when in the evening they said like, of course your model doesn't work.
And I go, what do you mean?
Your model doesn't work.
They're like, yeah, you know.
People ask us to do a thing.
They don't know how complicated it is.
So the specification, more or less, is incomplete.
So I know that the first thing I do when I get something like this, yes, I give a quote, but then I meet with them at the place.
And then 90% of the times I need to increase the price because the specification was not complete.
They did not see the complications.
I see that the place wasn't painted for 10 years.
In the kitchen, there's oil and fat on the walls everywhere.
I actually need to tear down the wallpaper.
So that's increasing price.
The first experience the customer has with me is a bad experience because I'm increasing prices.
That's bad customer experience.
So by principle, this platform does create a first bad customer experience.
I, as a craftsman, live out of...
good customer experiences.
That's why I don't need to do acquisition.
I get called.
But your platform is breaking that pattern.
I'm not getting called anymore by any of these clients.
So my customer acquisition cost in our terms is driving like crazy.
It's just doing a bad job.
That's why you see 30% of churn of craftsmen on that platform.
That's why you only see bad craftsmen on there because they are basically reaching for the strong.
The thing is that again, All of that is really simple to see and figure out for an agent.
It's hard to accept and to see for a human.
And that brings me to the last point.
A friend of mine built an AI-based platform which is creating an insane process which basically replaces all the research work and all that stuff and does it really deterministically.
You look at the results and you're like, Jesus Christ, that is...
Product work, state-of-the-art, par excellence, insane.
What I told my friend who built that platform, he's getting all these investors called.
I'm like, I don't think anybody will buy this.
And that's so interesting.
Now you have the dream product machine.
Yeah.
What's the problem?
The CEO is not hiring a perfect product machine.
He's hiring a human he can socially rely on.
He's hiring for loyalty.
If I'm having a bad time with my company, will you be in my boat and help me?
And we do whatever we can.
He's not looking for correctness.
And he will not believe that correctness.
What he's like, the next quarter is looking really bad.
Are you with me?
Are you going to help me?
That's what he's looking for.
He's not looking for the perfect answers.
If the PMO would come up with these perfect clinical answers.
The CEO wouldn't understand them because he, again, doesn't know how they are produced.
He's looking at the outcome and said, yeah, maybe.
So what he's relying on in the judgment, in that judgment part is like, do you believe in it?
Yes, I do.
Where did you get the idea from?
From AI.
I don't know.
Really?
What?
And that's so interesting.
I'm perfectly sure that software like this will not sell for the next three years.
Maybe then it will because like, and that's a huge part of the quiz.
It requires a certain culture to be able to rely on these things.
Because right now there's two aspects still in companies which don't allow us to rely on these results.
One is I don't want to be caught with using AI for my results.
That's why people are still painstakingly erasing M-dashes out of their end reports and all that stuff.
The other thing is ELSE.
If it's coming from you, it's AI slop.
If it's coming from me, it's good AI results.
Again, because I didn't witness what you were doing with AI.
I only see the result.
I can't see the reasoning, all that stuff.
I can't see the process.
But I think like the engineers, I've been interviewing engineers, right, for last year.
And they're faced with this because in engineering, like it's kind of similar to Do you review the code before you push to production or do you not review the code?
Right.
And I think there's two camps there are.
And I've heard a lot of engineers say, I review the code when I'm at work for my personal project.
No, but when I'm at work, yes.
Why?
Because I am accountable.
Yes.
If the thing breaks, it's on my ass, right?
I am accountable.
So therefore I review the code.
But there are some who are saying, and then there's like Anthropic who is saying, we have this distinction between leaf nodes and the other, the non-leaf nodes, right?
And for leaf nodes, they're like basically triaging, right?
They're not going to mess everything up.
So we're going to kind of be a little bit less careful there.
And then there's these engineers who are saying, I'm going to move it all to harness the guidelines, basically codifying the rules really well.
So that if I codify the rules super well, I don't have to review what comes out on the other end because I can trust it's good enough for specific things.
So I think engineers are already doing this process.
I think we inevitably will also need to do this process where we're like, okay, these things are kind of lower risk.
These things are higher risk.
So the things that are our leaf nodes or whatever they are, we just need to get really good at codifying the rules, which we can.
I'm starting with that in my skills.
I'm getting quite far.
right?
We need to codify the rules and then we need to trust that it's good enough, right?
So it's basically the exact same process where I think 80% of engineers are still at work.
I must review the code because otherwise I'm going to get in trouble where I think we're going to move away from that, where the expectation will be that you are just that good at codifying the rules that you don't have to do that anymore.
And the same should happen to product managers just there.
I want to ask you a different question, which is about, because we talked about deterministic checks and semantic checks.
And I feel like there's a lot in there because I feel like that's where the big difference is, right?
Like I find it harder to step away from my autonomous loop because so much of it is semantic.
When I'm thinking again of this thing of like plucking opportunities out of an interview, that's already like, okay, this is what the guy said, but I need to abstract that to some degree so that...
the things that other people said might also fit that umbrella, right?
Because otherwise I'm just going to have a huge list and that's not the point.
The whole point is to be able to cluster them.
And there's already something quite semantic in there.
It's like, okay, how do I take what somebody said and abstract that to the right degree that it still means the same thing, but now other things fit under it as well.
And then later I'm going to have all these opportunities and I'm going to cluster them because is what Marcus said the same thing as what Elsa said or is it a slightly different thing?
And I'm thinking about, okay, can I move these checks from semantic to deterministic?
Can I?
And also, should I?
Is that the point of the effort?
And I would love your insight on that.
And I, like, there's two things.
First of all, we are in the middle of that huge technical transformation.
And what I'm researching on right now is what is the difference between human in the loop and human on the loop?
Like, that's a very traditional technical problem.
terminology like you know red lights were invented not because we could trust red lights but because of economical reasons like in london it wasn't affordable to have a cop on every street corner so they basically come up with street lights which were first human operated and then the human got out of the loop on the loop so basically the red lights would be calling a human into the loop when they have an exception but other than that Red lights are just working.
And of course, like every cop standing on a corner, first of all, would say like, this is a deeply human activity.
Like no machine ever can regulate traffic as much and as well as I can do it.
Turns out a couple of years later that a red light is actually much better than a human because now we're actually breaking down when the red lights are not working and cop is standing on the crossroads.
And then we learned more over the years that actually not every crossing needs a red light, that, you know, circle traffic is also a very good solution for many things which are not that highly traffic.
So we learn along.
But I had the discussion yesterday at a conference.
A person came back to me and said, like, yeah, yeah, yeah, but we learned all that stuff about circle traffic.
We shouldn't have placed that many red lights.
Yeah, but in 1920, you didn't know that.
Yeah, exactly.
I get it.
And right now we're at the energetic engine.
Same thing, right?
It's like we made that decision based on the knowledge we had then.
And now we have new knowledge and we can adapt, right?
That's the whole...
There's the point.
Yeah.
So...
But engineers have it easier.
So basically, that's the step between, you know, in the Shapiro scale of coding, like the step to more productivity is between level two and three.
Level two is basically I still look at the code line by line, look what the thing did, and then I accept it, and then I go into the PR and do the manual PR.
Level three is I don't do line by line checks anymore.
I have all this test harness basically after the coding that checks for correctness.
And then if the machine says like, there's an exception, then I'm getting pulled into it and then I should look into it.
That's actually where the magic happens of productivity gains and everything.
We as product managers are not there because we don't have the test harness after creating the stuff.
And that's where we need to invest, as you said.
And again, like why I'm telling the red light example is like, We will not have a choice.
The cops didn't have a choice if they will operate the traffic anymore and regulate the traffic.
People said cops were too expensive.
People will come around and say like, you know what, product manager, you are just really a glorified, creative person that is actually just orchestrating insights into machines and then there will be results.
And the other fun part, like in the human world, in the loop versus on the loop thing is like dangerous example the same works for for drones in warfare like i'm really sorry to speak about that but drones are basically not human controlled anymore in the last mile they're on autopilot the the success rate of a drone is roughly 75 on the last mail on which it's autonomous which is not as good as humanly operated, honestly, but economically.
Yeah.
It's a dim one.
10% less effectiveness or efficiency of these drones doesn't matter because you reduce, and that's the actual number, human pilots from 540 in America to 170.
That's a huge number.
And now the thing is, As our barrier that we discussed in the beginning, a third to a third to a third is so low, getting us on the loop rather than being in the loop is not that high of a target because we're basically rolling the dice with our human judgment.
So, agents doing the same judgment on that same level will not be that hard.
So, some people on that planet will figure out that the jobs that we're doing as product managers is much more deterministic as we think, is much more collecting data.
orchestrating the data, putting it into the right channels, which right now is stakeholders and stakeholder meetings and back-to-back meetings, which are not necessary anymore.
Somebody will figure out that we should just context engineer and data engineer that stuff into the right pipelines and come to the right decisions with the machines and that's it.
But I don't think even, I'm not sure if it is that deterministic, right?
I really still feel like a lot of things are...
It doesn't have to be is the point.
Exactly.
That's exactly what I'm thinking.
Like, even the stuff that we think is semantic, we suck at it as humans or like, we're not amazing at it as humans.
So, and also exactly like if we have something, let's say the, like Marty Kagan doing the work is a hundred percent.
And then we can build an AI pipeline that does it at like 60% of what Marty Kagan or Teresa Torres would do.
And that's still like where most product people, product organizations are at like.
zero to two percent right because they're not doing it at all or they're doing it really clearly that is still like way better yeah i mean there's this uh there's two sides of it the first one is the very famous eric reese uh curve where he's basically talking about like um there's this curve about um how many insights do you take out of interviews and then there's this um it then it flattens at a certain point in time and he says like But most people are reading the whole curve totally wrong because the thing is on the left.
Like when you do not do any interviews and customer contact, you don't go out.
Like then there's no insight at all.
And that's basically the job most of us are doing.
We're sitting in our chambers and doing stuff basically based on our fantasy and imagination, whatever.
And that's the basis of how we now comes the hard part, how we do our job, which before was.
We are the filter of what needs to go into development because development is so expensive.
So I'm always talking about limited resources and I do a lot of overthinking because I'm the last filter.
Like we could do 15 things.
We can only afford one thing.
And I, as a human, with my limited knowledge and the limited data set, I'm trying to decide what is the one out of 15 things that goes into development.
based on all my biases and that's what we think is my in intrinsic human capability that i am the best tool to choose the one out of 15 things like like how insanely narcissistic can i be to think that that this is good judgment and then we come with all these um cost of delay things you know all the scientific stuff where we we don't know Any of the two numbers, the denominator or the other, we don't know how long it takes.
We don't know what the outcome will be because it's all guesses.
But based on that, we do the prioritization.
But now the world has changed.
Like if we have 15 ideas, we basically probably can have 10 of them go into development, then have a look and say like, hmm, this doesn't look.
like I imagined it to be like and my comparison is the switch between analog and digital photography like ages ago what you would do is like you would buy an analog film 24 to 36 exposures cool and then because it's such an awkward process like after your vacation you go back home you have all the films and you send them into a lab and I don't know if you're old enough to remember that but then you wait for weeks and then you get if you're not spent all the money, you get the contacts of these things and you choose the ones which you want to really amplify and, you know, make bigger and have extra paper prints or you have all of them printed on really small paper.
Anyways, but hopefully out of these, I don't know, maybe a hundred exposures you would make over your vacations, hopefully you show maybe 10 to your relatives to show how great your vacation was.
Now with digital photography, you don't think about taking any photo anymore.
And probably over the same vacation, now you take 500 to 1,000 pictures.
But also, hopefully, you will not develop any of them on paper, maybe one.
And to your relatives, you also only show five of them and not 100 because you made these many pictures.
And it's the same what we have to do.
We can look at all of them.
of the things, the 10 features, we can have them built in 20 minutes and iterate on them for another two days and then pick and choose ever more and say like, this is totally going into the wrong direction.
More of this, less of that.
And then only again, after two weeks or whatever the cycles are, hopefully only one feature survives out of the 10 we tried.
But now the judgment comes after the fact.
Yeah.
Because we didn't have to filter based on our bias, but say like, yeah, let's be open.
Let's look at the stuff.
And that's what actually the journeys are doing, the Boris journeys.
They're like, oh, we could do this.
We could do that.
And then they're having all these terminals open, but they're not shipping all these features.
They're basically saying like, that's crap.
That's crap.
That's crap.
Like, how stupid could I be?
to even try that out.
Let's scrap it.
And that's your new job about it.
And what it means is we have to be really good at opening up, looking at the thing, discussing with the team over and over again, meet with the team probably five times a day, not once a week or twice a week or whatever, and then have total filtering against what is CI, like what is even integrated into the trunk.
What is the release candidate?
What then gets released to the client, actually?
Because the hardest loop we cannot improve and fix and accelerate is the maximal acceptance of features in the market.
Because just because we can build the market.
And there's another example I like about this, like, again, eBay.
So we told people for 10 years that auctions are the best thing on the planet.
Like, whatever you want to sell, put it into an auction.
And then, of course, Amazon came along the corner and they're like, hey, just buy stuff.
Fucking hell, people can just buy stuff?
That's insane.
But people actually like it.
So after educating people that auctions are so good, we're like, yeah, maybe, you know, buying stuff on the fly is also a good idea.
And it took us five years to convince people that.
actually, you could buy stuff at eBay.
Also, we sucked a little at it and the logistics went there and all that stuff.
So anyways, we had too much of this heritage of auctions and, you know, it was a mess, basically.
It took us five years.
But here's the thing.
The customers wouldn't have, there wouldn't be any way where more and sped up coding would have helped us to speed up those five years.
No amount of code being produced would have helped us to get more acceptance in the market earlier for that new idea.
So that layer is still there.
So while it's nice to do more and to try out more, we should actually use it against that human inability that we cannot imagine things that well and we don't have an idea of the future.
What we're really good at is looking at things and say like, oh, now I get it.
And that might work.
Rather than having the fancy idea of I make up a concept, I think a word that only exists in Germany or Europe, conceptor.
I make a concept of a feature out of thin air.
I try to convince a bunch of people of it and that's the best future.
Again, future that might be out there.
What?
Rather than now.
Looking at 10 futures right now, making it a present, again, idealized design, looking at that present day thing, how does my product look like with that feature in it?
Oh, that actually looks good.
Let's try it out in the market, but let's scrap the other nine things.
The two points there, I think one thing that you found that I super agree with, but I think we should highlight, like just because we can produce a lot of versions of something really fast.
Please, for the love of God, let's not put all of them into the market.
Please.
Because we're dying here, right?
With so much shit.
And it was always a problem.
And with more marketing and more sales and more products.
And now that's even worse.
So please, for the love of God, not do that.
At the same time, now you're inviting taste and judgment back into the door, right?
Or we call it discernment.
It's also a term that I heard, which is basically now I can quickly look at 10 things.
And I need to pick which ones are worth progressing on.
And I'm not going to collect real user feedback on 10 things because poor users, like, let's please not do that.
So now we're actually going away from evidence-based, right?
And we're going back to taste and judgment, which hopefully this taste and judgment has been built on the back of years and years of actually seeing what works in a market.
So our taste and judgment is somewhat evidence-based.
It's formed through...
following due process over like decades of work hopefully or it's literally just going yeah i think that looks cool i think that doesn't look cool so right like because now we are inviting taste and judgment back in through the back door to saying okay these ones i'm going to put into the market and these ones i'm going to toss into the bin absolutely but i want to make one difference explicit and like uh how i try to express it is i think it's a difference between buying a car based on the marketing or having a test drive with it actually because basically by building the feature based on my fantasy and my concept or whatever you want to call it my prd whatever that's basically committing to a thing based on on the marketing like the the description of what it might be but when we judge the thing based on it's already built we see it in context of the product that's like a test drive that it doesn't tell me how the the car feels in five years and if it actually was the right choice there wouldn't be a better alternative but it gives me a very good impression between five cars so i said out of the five cars i had the chance to test drive this felt best Yeah.
And that's what we're really confident with doing decisions on.
Whereas the other one is like inviting all the sunk cost syndrome, like now that we spent four weeks already on the thing, let's go the last mile, let's build it, let's push it to production.
And then finally, when we went that far, let's push it out to the customer if he wants it or not, who gives a shit?
And I think that's a totally conceptually inherently different thing, judging it in the whole context, again, in the sense of creating a better today.
rather than imagining things and then pushing them through and i think that's totally important and i think i could live with that as a current state like do the job like let's figure out the 15 things which look very probable based on the semantic checks oh and i wanted to say another thing on that one which i will still do based on on our correctness and completeness checks we did that that's Very highly probable 15 PRDs for 15 features.
Let's give them a try.
I'm not sure.
I'm still an idiot.
But these are my best guesses based on all my checks I can provide right now.
These are all candidates.
Let's have a look.
And then as a team, maybe, or whoever does the decision, based on the real world, let's make a decision.
What I wanted to come back with the semantic checks and the vagueness or not is...
If we think that a thing is not to be decided right now, we don't have complete correctness and all that stuff.
I think another model is basically statistical forecasting.
What is the probability of that being correct?
And I think that's something you can build in your skills.
How much do you think this is validated already?
give it a percentage and and again machines are really good at that like the machines do know that the information is partial that they have and they're like really good at giving numbers at that so if you give a thing a probability range like even a range like you can if the thing then says like this seems to be a valid feature for me 50 to 80 percent you see like oh yeah that's still a 30 percent bracket that that's still vague 50% is not really high.
80% is not really close to 100%.
So if another feature comes along in your same run, which says like, this is 70 to 90%, I know what I would much more probably pick and which I would set into the race.
And I think that's a very valid probabilistic statistical approach for a lot of things.
I feel like, you know, health does the same thing.
You go to a health check and most of the time they're not saying the reason is that.
What they're saying is the probability that it's this is 80%.
And then you're like, huh, 80% is not a lot.
But then what they tell you is like, in health, 80% actually is a lot of safety already.
And I think like when I think about my knowledge from building LLM-based workflows at the company where I am now, where we're building this thing.
which is also quite semantic, where you basically have an email coming in and you build a workflow where you equip an LLM to take this whole email and we have like 130 case type labels and it has to associate one or more case types to it.
So it says, basically, this is what this email is about.
And that's very similar to this semantic check with what I'm doing, as in like, here's somebody saying a bunch of stuff and pull out opportunities.
And there you think in these...
like workflows right instead of just going what my skill is doing is saying like here's a transfer pull out an opportunity but if you have an agentic workflow you say you break into little pieces right you're like making it smaller and then you might have and now I have some rag system and I'll have some examples of things where it went really well and they are similar and then maybe I'll have a next step where I'll pull out some anti-patterns of like, this is what you usually get wrong here.
Don't fuck it up like this.
And then you'll have some confidence routing, right?
You'll be like, and this is maybe your probabilistic calculation or something like, like how sure are you now that this is correct?
And then you'll have some human in loop workflow where if it is below a certain star, you'll have a human in loop.
And if it's above, you might look it through.
So if I am rethinking my skills as an agentic workflow, it would be something like that, right?
Where you do have these.
confidence scores and probabilistic measures bakes into it at the end which my skill isn't doing but i would do that if it were an agent and that again funny enough is coming back to what i said about human in the loop versus human on the loop like in hindsight there's always judgment about when does human on the loop make a lot of sense and like the decision criteria is like is there existential harm lethal harm involved potentially as a risk.
Probably for moral reasons, you then want a human in the loop.
Just for moral reasons, you need somebody to be guilty, to scream at and so on.
And we know that most of the times, especially in medicine, the system comprised of a human and the machine is better, even if the machine statistically is better than the human.
The combination of the human with the machine is always, most of the times, providing the even better results.
Then another thing is reaction time.
There's processes where the human in the loop doesn't make any sense.
That's where, for example, aviation comes in.
We know that automated landings are much better.
So the human is only called into the loop in an enormous escalation.
When we think the human in the loop is better than the fly-by-wire system of aviation.
And now we're coming back to as obviously and objectively as humans we suck at product management given the numbers we have.
Again, the risk of doing experiments, people like you and me, with human just on the loop, being called into the loop.
oh, I have ambigdata, can you clear it up for me?
Have you more data sources which you can supply to me also, like if I'm the agent?
I need a decision here being done because for me it's flipping the coin.
I don't want to flip the coin as an agent.
These are, I think, totally valid experiments.
I'm pretty sure with just a little investment, objectively we get to better results than the human on their own.
If, again, The CPO from the PO wants a person to be screened, I'd find, have a mixed system.
But I think most of the work can be done without human intervention, just coming up with the choices.
Yeah.
And so just one, two last thoughts.
Sure.
One is for me, like, so I built my pipeline of discovery skills to do these steps, right?
And then I've been toying with them and having some fun.
And then I was like, how about, I'm not going to look at the output at all.
Mind you, it's not an agentic workflow.
There's no rag.
It's really just a skill that has my rules.
And I have this set of data.
And I just ran it five times in a row, every step, without looking.
You know, I'm just like, okay, take the output of the previous one, keep going.
And I was like, let's see if it picks the same number one opportunity five times in a row.
And it did form a different site differently, right?
Because it's always abstracting, but it's the same one.
And I was like, that's insane, right?
I was surprised by that result because I think, and I'm also not sure if it's good.
I think it's good.
But on the other hand, if humans do it, it's slightly different.
And maybe there's also something about diverse viewpoints and because it's literally the same LLM doing all this.
But I was like, wow, that's pretty consistent for something that is semantically very heavy, which I found a very interesting finding.
Yeah.
Yeah.
You said you have two thoughts.
Yeah, exactly.
So the other thought is based on what I told you in the beginning, but I just wanted to say it, is my bubble is screaming about delivery is solved.
Discovery is the bottleneck, right?
And it's this idea of if this is the bottleneck, right?
We're kind of saying that this is the problem that we should address, right?
A bottleneck, you don't want a bottleneck, right?
It's a problem then if we say that.
But then when we say, okay, so let's solve it, right?
And we also know that...
we're as humans quite poor at it and we have all of the data to back that up as you've said many times so I'm like okay let's address it and let's try to figure out how we can have an autonomous agent basically doing this work for us or at least greatly assisting us but then we're interjecting going taste judgments moats blah right and this is so funny because these are two completely conflicting thoughts I think to be having yet we're completely comfortable in saying them out loud yeah and I I think that has to do with the future being so drastically, quickly being so unevenly distributed between different bubbles.
So what I said before, this strange system, which is taking care of the business case and the technicalities behind it and runs for one and a half, two hours.
I guarantee you, and I bet my right hand, that I have not met a person who injected their problem into that thing and not came out totally astonished about the results.
the angles it's opening up and they're like pro bono I'm coaching a person with a private business and I sent his problem into that thing and I sent him the result and he looks into it and he's like Jesus Christ this is mind boggling I will take the next half year just to cover 20% of the ground that thing is opening up for me for my business and I believe 90% of the things that it claims I just need to execute based on that thing right now.
It's so crystal clear.
And I don't know why I did not have those ideas or whatever.
So I'm not a genius.
I built that thing.
It runs for two hours.
It comes up with these results.
What you see out there is the people doubting that stuff.
Totally legit, totally fair.
But the reason they doubted it, they have never seen these results.
They have never injected them into that world.
And now it sounds...
totally stupid, unfair, a promise for the future and, you know, all that stuff.
But all I say is whoever is exposed to that techniques, who goes, for example, through my course, and I don't mean it as marketing for my courses, the people who see the results of context engineering techniques in cloud code, codex or whatever.
It's a mind-boggling aha effect where they're like, the quality is not the problem.
The explainability is the problem.
The explainability to other humans, basically.
Because the results are here.
And again, the bar is so low, it's not hard to be better than me.
Totally not.
And then creating the social fabric again around this might be the art that we need or whatever.
But the sheer results are really not the bottleneck.
Especially in research and discovery.
All that I'm talking about with that two hours running agent is basically discovery.
New angles in my business and so on.
I think for those who look at the results, it is a solved problem already.
I think discovery is a solved problem with LLM.
Can it be better solved still?
Sure, no problem.
In five years, also the same thing.
company operating systems where they open up something like cloud code for an enterprise in the morning and that's the operating system of the company that's going to be solved problem in five years and and what's going to happen is that that the bubbles that exist right now who saw the results and the other ones who like none of this is possible will merge more and then it will be accepted and Right now, it looks like machine storming and everything, and we're too positive, and it all is one large bubble of all the admittedly societal, economical, ecological unsolved problems with the technology.
I get it.
But all of that is put into one big lump of feelings about this can't end well.
protects a lot of people from actually accepting the results we have with these machines.
And I would encourage people, like, if we have all these negative side effects, I get it.
Yeah.
We will probably not solve it in a year, but it is, like, again, like with electrification, like with motors, there is going to be a societal solution around these things.
There will be losers.
There will be winners.
I get it.
Like with cars, like who would imagine the infrastructure we have for cars right now, a hundred years ago, does it have negative side effects like crazy?
Would we wish a world in hindsight without cars?
Probably not.
I don't know.
Like, I think that's very subjective.
Same thing with LLMs.
Will they go away?
No.
Like, even if all the entropics and open airs go broke, the technology is out there.
The Chinese will do it their way.
The LLMs will be out there.
People will adopt it because for economy.
It's just an economic fact of the world that we cannot ignore it.
And I urge people to lean into it and learn the positive effects and not just to rely on the negative side effects and benefit from it as a species, basically.
I think we could do a whole other recording on also acknowledging the negative sides.
Yeah, totally.
We sound like...
We're totally hyped and that's absolutely not how I self-identify at all.
Especially at the moment and probably, as you say, with cars forever, right?
There are also very negative side effects that I can talk about for many, many hours and things that I find wildly problematic.
But that's not the point here.
Yeah, I have to go.
I have to pick up my ticket from the sleepover.
Yeah, and that's important.
fact and part of human life.
So take care of that.
And thanks for the conversation.
Yeah, let's have another one on the other aspects as well.
And maybe we also got further on our research and stuff and all these things.
And thanks for your time.
Yeah, see you soon again.
And have a great day with picking up your kid and then do your work also and all that stuff.
So have a great time.
Thanks a lot.
Thank you.
See you.
Thank you.
For listening, Stories Connecting Dots, a business podcast to discover a world full of possibilities through stories told.
Presented to you, Markus Andretak.
