# Controlled Agents For Enterprise Data Security

**Podcast:** The AI Native Dev - from Copilot today to AI Native Software Development tomorrow
**Published:** 2026-07-21

## Transcript

OpenClock could really do everything.
It felt very uncontrolled and just a disaster waiting to happen.
LLMs in an agent harness, if they have an open check to do everything, then it's inevitable that they will eventually do something terrible.
It's just like evolution.
We came up with this idea of having an agent that's within the controlled environment and has access to the data.
And you can have it investigate the issue and...
troubleshoot it and it has access to everything, right?
But one big constraint, it cannot output except for a structured validated output.
So it can only output for a fixed vocabulary, which you can verify doesn't look any data.
So those two agents, they're called Mulder and Scully because they worked on the X-Files.
The AI Native Dev is a podcast for developers and engineering leads at the cutting edge of AI and agentic coding.
Join your hosts, Guy Pajani, and me, Simon Maple, every week as we chat with the most exciting voices in AI and tackle the biggest questions facing developers today.
This is the AI Native Dev.
We just wrapped up two amazing days at AI DevCon in London.
But the great thing is that we get to do it all over again in New York City this November.
You're absolutely right.
We're going to be back in the city that never sleeps on November 3rd and 4th for more amazing sessions.
really engaging hands-on workshops and much more.
Yep, all that great networking, partying, eating and drinking that you've come to expect from AI DevCon.
We think we have one of the best hallway tracks in the business and it's the perfect complement to our incredible speakers and presenters.
We'll both be in person and virtual with live-streamed access to all main stage keynotes and talks.
Sign up right now for our Super Blind Bird ticket for just $100, only available for a limited time.
We're really excited to be headed back to the Big Apple.
We hope to see you all there.
Hello, everyone.
Welcome back to the AI Native Dev.
Today, we're going to have kind of a real-world case study talking about building agents within a company.
This will be within Sierra, which is a data security company.
We'll hear more about that in a sec.
And to talk about that, we have Ori Shoshan, who has been one of the key leads building that out.
So welcome to the show, Ori.
Thank you.
Nice to be here.
So we're going to go a little bit in chronological order.
So just to give people a little bit of a teaser around what it is that we're going to talk about, I think you've built a really, really interesting and oftentimes quite unusual platform for how to kind of work with agents.
That's a compliment.
It's like you're a little bit weird.
It's good.
Who wants to be normal?
No, actually, like...
A lot of many smart sort of small innovations and learnings.
I like you've built a different approach to harness engineering that starts maybe like whitelists tools versus blacklists them.
You have made some really smart infra investments in creating context graphs.
It's my German side that has to have control over everything.
It's just that sort of efficiency and path on it.
But that's been great.
And then, you know, a lot of things.
I love the sort of small learnings around how.
Do you permeate through the org?
You know, things like having people take ownership of their agents or they care about it.
So there's all these like great nuggets that we're going to dig into.
But instead of jumping into those, let's do this properly and kind of run through a little bit of intros and origin stories.
So first, tell us a little bit about what is it that you do at Sayara and maybe a little bit about what Sayara is.
I'm a tech lead at Sayara.
I work all across the org.
A lot of my work is focused on the DSPM product, so that's the data security posture management, as well as leading AI adoption.
Sayara is an AI data security company.
We combine AI security.
Well, everybody wants to do AI security nowadays.
And before AI really boomed, Sayara did that data security, right?
So we're a data security platform, meaning in a nutshell, we scan your data.
We find where the sensitive data is, what type of data it is.
And then we let you do the really important stuff with it, like deciding what your posture should be, seeing what's exposed and how it's exposed and fixing that.
And in AI security, also enabling you to...
allow AI access to that data, but in a controlled and secure way.
Which is, it's quite important to know where your sensitive data is and what you do with it if you want to build agents and let agents run freely.
Because otherwise your risk leaks, and if you're scared of data leaking, you're kind of paralyzed, right?
So it's really all about control.
Like to my German side.
I love the generally analogies in companies that have kind of grown up.
out of the problem of the empowered workforce, which I think CERA and sort of data security poster management or sort of DSPM, you know, fits that sort of mold where everybody gets to do a thing, like everybody gets to create data in the cloud, you create those, and then, of course, you have chaos because everybody's empowered to do that.
And so here, like, income, a variety of solutions, many in the security space, specifically in DSPM, to, like, wrangle where do I have data, what's sensitive and what's not.
And then as we get into agent land, you know, agents are even less trusted, you know.
empowered entities that you're sort of running along.
And so they have to be reined in even more.
And so I love these kind of realities in which kind of solutions built for the pre-AI era, but for the empowerment era, actually very much sort of applied to this.
Yeah.
Oh, and about me.
I forgot to tell about myself.
So I joined CERA about a year ago for the authorized acquisition.
So Authorize was my company.
I joined together with my team and my co-founder, Tom Greenwald.
And together we lead AI adoption in CERA R&D.
And that's sort of how it all connects.
And I guess just for context a little bit, that's how I got to know you.
So I was a part of the Authorize journey.
It was super interesting to talk about sort of intent-based access control and all that.
I still believe in that.
Me too.
It'll still happen.
But I love sort of working now on the CERA side, collaborating there.
So this is sort of the background.
You're sort of in, say, a clearly data security company, I think in the order of magnitude of sort of hundreds of engineers on it.
How did you come?
So you're doing AI enablement or you're building it, but I think your origin story was more about like scratching your own niche, right?
Right, so it's 2026.
And I'm sitting in my living room.
It's January.
So it's now 2026.
We're still in 2026.
I don't know.
Timelines are a little bit murky on the island.
But it's January.
So we're like five, six months.
It's January 2026.
And I'm sitting in my living room.
And I'm doing something on my phone.
And I ask myself, why am I using my hands to do this?
Why is it not an agent doing it for me?
It's 2026.
It's not 2025.
As we all know.
Late November, December.
It's like a cave in love yourself.
Yeah.
2025, Frontier models really took a leap, at least from my perspective.
I saw that I was able to build a primitive harness back then, basically, you know, like a coding agent with some tests and allow an agent to run much longer than it would have been able to before with just one generation of Frontier models before.
And then, so a few weeks after that, I'm in my living room thinking, why am I still doing things on my own?
I need a personal assistant agent to do it for me.
I'm building a house and all that, and I needed something to keep track of all of my commitments and the argument I currently have with the contractor for the house.
And I really wanted an agent that would...
kind of be like OpenClaw, but I didn't like OpenClaw at the time.
I think it was still called its previous name, CloudBot or something.
MoldBot, CloudBot, one of those.
Yeah, it went through phases.
It's like a teenager growing up.
But I didn't like that OpenClaw could really do everything and anything.
It felt very uncontrolled and just a disaster waiting to happen.
Because I kind of view AI as, or rather, not AI, but...
LLMs in an agent harness, if they have an open check to do everything, then I feel like it's inevitable that they will eventually do something terrible.
It's just like evolution.
They will explore that path if it's there.
So I felt like hooking up OpenClaw to all my stuff and having it manage my life is kind of precarious.
So I started building out my own personal assistant, which...
where I took a bit of a different approach, where instead of letting it do anything, I let it use a specific set of tools.
And then I got it on WhatsApp, and I would text with it, and it would be able to mention where its data is from.
And then I figured, actually, I would like to know it's not making stuff up, and where is that sourced from.
And I figured, well, I actually have that data.
I have that link.
So why don't I just have the harness tell me what the source actually is, right?
And then eventually I realized, actually, I want this at work too.
I mean, why am I still reading Slack messages and following on Salesforce?
Most time consuming tasks.
Yeah, well, why don't I just know everything all the time, but only the most important stuff?
And I want to have links to all of them.
I don't want to wonder, is this something the model made up or is this backed by data?
Yeah.
So this is, I want to sort of unravel a few things.
I love the sort of the personal to company paths on it.
I think in part because you're kind of allowed to play a little bit more at home, right?
As much as you are cautious about your sort of your own kind of personal blast radius, you know, it's a little bit different than playing with sort of company data.
So as you build those out, I think the first thing you talked about is sort of this sort of containment of...
of only giving the agent some specific subset of tools on it.
I guess how crippling is that, right?
Because some clearly come from an AI or sort of AGI maximalist path.
It's a lot of like, you know, do not get in the way of intelligence, right?
Just sort of let it do the path.
I guess how much was that felt that, you know, if you contain it and you need to choose every time which tools are available and maybe like a little bit of...
There's probably still some leeway that you have to give it.
So what did you find was effective in terms of the lines?
So I would give it quite a bit of tools, and it would be the one tool I would not give it is Bash, which is basically do whatever you want.
But it would have a bunch of other tools, for example, to read all my WhatsApp messages and browse my email.
But how do you say the attack surface?
uh for each of them would be well known um and i'd always leave an escape hatch so i could have uh so what you don't want is for an llm to try and use a set of tools and really be trying to do something else if it had bash let it do whatever and then do something that doesn't make sense just because it doesn't have the option that's a bad failure mode because you may not notice unless you actually look at the thinking tokens for the model.
So I always leave a tool that's a sort of escape hatch, whether that's doing something else but with a human in the loop or recording feedback.
So for the model to be able to tell you, I wanted to do this, but I have no option to do that.
And then you go and add that.
And you very quickly arrive at a set of 50-ish tools.
which the model has access to, and it can do pretty much everything it wants.
But you know that it can only do those things and structurally, or I guess, de-harness and force us that it can't do anything disastrous.
Like it can't send all my personal notes to the contractor over email.
Yeah, that sounds like definitely a negotiation sort of side, or maybe even just sort of a goodwill.
So, okay, let's sort of...
Come back to that a little bit in the context of the company on it.
The other piece you talked about is this notion of citations and sort of bring proof.
I guess it's sort of the association that comes to mind is what we see in Google, right?
Oftentimes today, when you Google something, you get the preview.
You've got sort of the equivalent of the blue links today, right, that talks about those references.
Is it just about sort of sourcing the information?
Is it similar, you know, what you built for the personal agent there?
Or is it something a bit more?
distinct from that when you use it for your sort of maybe kind of more elaborate because Google is very much about like lookup of information so maybe a little bit more expected that there will be a link for it.
Yeah.
So there is I took inspiration from ChatGPT actually where they at some point they started adding links next to where the claim is sourced and But the core need I had was to know if the model is lying to me easily because you get like a wall of text from an LLM and then you're thinking, okay, well, which of this is based on fact and which of this is reasoning that's correct and which of this is reasoning that's incorrect.
And it's really cognitively exhausting to have to constantly evaluate that.
And it's one thing when you're working with a coding agent and you're in an interactive session and you're in the zone.
But for the personal assistant use case, you really don't want to be paying so much attention because that's kind of the use case.
So I wanted...
So you don't care about, like, you're fine for it to take more turns for the exchange of getting the correct result at the end of it because you're not sitting there supervising.
But if you get the wrong answer or you get something, that can cause substantial damage or at least waste a lot of time.
Yeah.
Well, I sort of assumed that...
LLM is going to get, the model is going to get some stuff wrong.
And then I wanted to make it more obvious to me when that's happening.
And the way to do that is to have, basically, instead of having the model be able to just produce prose, it would be able to produce structured output.
And that output is a set of blocks where each block is a tuple of a citation and a source reference.
sorry, a claim and a source reference.
The source reference is a citation.
And then when the output is rendered to me, I just get a list of, well, each claim is just a paragraph, say, and then it has a source reference and the link is automatically rendered.
And once I had that, I realized that I could actually use that to do some context engineering and do a claim verification as well.
Right?
Because...
LLMs tend to hallucinate more or make the wrong conclusion when they have a ton of stuff in the context and they end up paying attention to the...
They're kind of in the dumb mode.
A lot of their context window is taken up by historical actions.
They might end up answering the wrong answer because they're kind of looking for a needle in a haystack.
They might look at the wrong needle.
Attention is the scarce resource, right?
Yeah, attention is all you need.
So I realized, okay, now I have the ability to take each claim and pair it with the raw data that's supposedly behind it, at least the LLM says so.
And then I can take those two pieces of data and ask another model with clean context, which has the claim and the data, is this claim supported by this data?
And in that mode of operation, they tend to hallucinate a lot less because there's no noise.
So that reduces the amount of hallucinations overall.
So each claim independently tends to be more accurate, that there's far fewer hallucinations.
And then if it is still a hallucination, I can just see the source.
And it's easy to tell if it made a citation that makes no sense.
It cited something completely irrelevant.
It does happen quite rarely because the LLMs, the models have gotten quite good, right?
I love the, so I like that.
And I think in general, kind of, I guess, a principle of working with multiple agents is that any individual agent might hallucinate.
But if you layer on even one more to sort of, you know, review the work, then you reduce dramatically the sort of the chances of a hallucination.
And I think you took it a step further because you're even requesting the original kind of claimer to create something in a slightly structured format on it.
So I like that.
So the escape hatch here was I let them actually go past it.
If they don't have a source reference, they can still produce a claim, but it has to be marked as opinion.
And that's where all the reasoning goes.
So I know to take those with a...
grain of salt, and the rest is what I use to evaluate the response.
I really like that.
I think oftentimes the thing that is annoying when you're sort of interacting, and indeed the LLMs have gotten better in this, but what used to be very annoying is it's not so much that the LLM makes a mistake, but that it says these mistakes with such confidence on it.
So even just labeling it as like, hey, I have proof of this, and I do not.
This is my opinion.
It's probably useful.
I think it possibly even helps the LLM.
to have that in the context as the output.
Right.
So it knows how much to build on it.
Yeah.
So let's sort of move this back into work.
So you've sort of learned, and I imagine you sort of started bringing this into the company before all of these things were sort of resolved on it, but I don't know.
But as you brought it into the company, so how does this manifest in bringing it into the company?
What did you have to set up, and what does it look like today when you run these agents inside the context of Sayara?
Running it at home.
I ran it on a Nuke PC, and I didn't even know it had a fan.
But it turns out it has a fan.
And then bringing it into Sierra, at first I started running it locally, but I wanted it to run, obviously, 24-7.
And I wanted it to have access to some things and be unable to access some things, which maybe my laptop can access.
Again, my gentleman has to know, like, what?
I have to know what's the possible failure mode and be okay with it.
Generally, those are good tendencies when you work at a data security company.
Right.
So after you set up the kind of the, so you handle the sort of the ops part of it, you're going to run this.
Now you have it in some sort of Sierra owned or sort of controlled agentic environment.
And then you mentioned that you have different agents that can run on it.
So that's already like a bit of a distinction.
So before you kind of created this one agent to be your personal assistant, you didn't need it.
probably like multiples of them.
I guess what, I'm curious a little bit because I think a lot of companies talk about like the blank slate problem, right?
Or how do they get going?
What's a type of agent that got created?
We'll come back maybe a little bit to how you create agents on it.
But maybe just sort of talk a little bit about the use cases.
So within the context of the company, How did you interact with yourself, I guess, and with sort of users?
Because at the beginning, it was like, you know, to help you read Slack messages.
And what are the use cases that end up, you know, proving to be valuable?
Right.
So the first one was a type of personal assistant or its work mode is chief of staff.
And then I looked at making it more generally available.
which initially the obvious choice was to make it into a Slack bot that you can DM and tag in threads.
A natural sort of chat interface for you.
Yeah, just bring the bot to where people walk already.
And so the first use cases, the easiest ones to build were an on-call triage agent where an alert pops up, there's a discussion about it.
You can ask the agent, hey, can you root cause this?
What's the impact?
Which tenants does it impact?
Stuff like that.
Very operational RCA type of agent.
And this is you create the agent or someone creates an agent, but then multiple people use it.
Like this is a deviation from what you had before.
Before you were like there was a one-to-one, one user to this agent.
Now there are multiple users to this agent.
But collaboratively or just like every individual might sort of...
at the bot and get a response?
So early on, really it was just every individual, but there were some emerging properties out of that.
Well, the agent can read the thread, so it can tell when, it can see the previous discussion, it can see people talking about using the agent, people complaining about hallucinations, and then when you tag it, it sees the prompt, but it also sees the context, so you might respond to something else as well.
Okay.
Slack is like naturally a multi-person, kind of a multiplayer environment on it.
So it just sort of, we got that for free.
So that was the interactive use case.
And then another use case was something our CTO Tamar proposed.
We have these channels on Slack where SEs, AEs, and so on might ask questions of product and R&D.
And you can...
answering some of these questions can be done by looking at the knowledge base, the internal knowledge base, the public knowledge base, sometimes the code, sometimes historical questions.
And I think in many cases, you just don't know exactly what you want to search for, kind of like searching on Google.
If you know the right keyword, you're going to find the answer, but sometimes you first have to figure out what it is.
Agents are quite good at that.
Yeah.
But in this case, like the way that people interact with it, unlike the previous one, is they don't add the bot.
They just ask the question in the channel.
And then the bot would ask them using an ephemeral message, which is the type of message that shows up on Slack when only you can see this.
Like when you tag somebody in a channel and Slack bot says, hey, this person isn't in the channel.
Should I add them?
That's an ephemeral message.
So the bot would ask you, do you want me to try and answer this?
And then if you click it, yes, it's going to answer publicly in the thread for everybody to see.
And that's actually an adaptation because initially it would just answer and it would annoy some people.
So I had it just start asking.
I figure just treat it like a human, right?
If you ask for permission and somebody wants you to, it's fine.
They won't be upset.
Yeah.
Cool.
Okay, so we kind of scattered a little bit over here as I kind of passed different paths.
But you have the bot.
It's running in your Slack.
It has one use case, which is when you run mode of operation, I guess, when you sort of interact with it.
And it sort of answers questions.
And one that it's sort of proactively, I guess, on a schedule or just monitoring a channel.
So it reviews every message.
So it's just monitoring the channel.
It's called ambient mode.
And it has a classifier that's...
uses a cheap model to decide, is this a question, is this an announcement or something else?
And only if it's a question does it actually offer to answer.
And there are a bunch of other use cases.
That's where the multi-bot part comes into play.
So initially there was just one Slack bot, my own.
I called it Borg for the Star Trek reference.
That's the cyborg hive mind.
And the thing about the platform is all the...
would-be agents of the platform are going to share the same knowledge.
They're just all going to be the Borg.
Yes, they're going to all be assimilated.
So the joke was too good.
Plus, I took Sayera's...
If you go to the website, you'll see we have this flying space pod icon.
So I put a little cute face on it, and that's Borg.
So then the multi-bots, right?
So I started out with just my own bot, but the way that you would interact with these agents is by tagging them, right?
So now somebody else wants to create a bot.
The way for them to do that would be to create a new Slack app and give it a name and possibly a personality.
And this was something I did for a functional reason.
I just wanted them to be able to customize the system prompt, the tooling, the context.
The platform is aimed at engineers, right?
So they would potentially rewrite the code that sets up the agent when it gets a prompt and be able to do context engineering.
But a really nice side effect of that was...
they ended up naming their agent and really feeling like it's their thing.
And it really helps adoption because it was no longer, you know, I'm adding a bit of code to Ori's agent.
It's my agent, which was really kind of my goal with this, for this to be a platform that everybody on CERA can use to build agents.
And they get the data infrastructure, the citations, the knowledge graph, which maybe we'll talk about in a sec, but also the guardrails.
and to satisfy my need for those controls.
So let's, cool, okay.
So at this point, we've advanced down the story of it.
Now we have these bots running.
They're connected to Slack.
They can do a variety of use cases, which we, you know, we touched a couple.
We probably should sort of cover some more.
But then they don't, you don't just sort of lump on all the many use cases just to Ori's Borg.
But rather now there might be, you know, there are multiple agents.
I love that sort of human insight of it, which is once you name it, you care for it.
It's fun, too.
They make stupid jokes while they're answering.
Well, and you can do because you can kind of add your personality.
But then if you care for it, there's just like the investment within that.
But it also lowers the friction because people allow themselves to modify it because it's fine.
It's mine.
I'm not getting in anybody's way.
So on one hand, you kind of get that.
foundation.
Maybe let's indeed go back a little bit and understand first, like maybe let's talk a little bit more about some use cases.
And then we'll come back to understand how did that citation platform kind of manifest?
Probably should talk about the knowledge graph there.
But how did that sort of come back?
But first, let's come back a little bit to the to the use case example.
So like one, I get its internal knowledge base look up.
I think everybody relates to that.
That sort of.
common use case.
You start talking about some ops issues, like root cause analysis of it.
Maybe say a few words about that, and I know you have a few other ops use cases.
Right, so CERA is a complex system, and by its very nature, data security deals with a lot of data.
Some of it takes quite a bit of time to scan, just because of the sheer scale.
And so some operations can run for quite a while, and that makes them hard to troubleshoot.
I mean, if something ends in 10 seconds, you...
You look at the operation.
It's done.
So these on-call or operational troubleshooting agents would have different triggers.
So some of them interactive.
Some of them are like the ambient answering questions agents, but they would actually respond to alerts or tickets from customers being opened and do initial triage.
Some of them would do triage in terms of...
assigning it to the right team based on what's in the ticket.
And then there's an entirely different class of agents, still all Slack bots.
But we have, say, the costs agent, which analyzes the cloud costs for Sarah and then creates a report and root causes changes.
So it would see an increase in cost and try to attribute it to increasing workload or changing the code.
and help you sort of keep it under control over time.
Kind of like spring cleaning type housekeeping agents.
Exactly.
And then you also have agents which help product managers.
For example, by looking at feature requests by customers or mentions on Slack for features by customers and sort of help the PM see that there was a trend forming way before it would have been detected otherwise.
Because before you had to have say an SE or an AE, say we need this feature, and then people would need to realize it's actually the same thing.
LLMs are good at finding those similarities and saying, hey, I think this is actually similar.
Pay attention here and here.
And this is like, how many people have built agents?
Because these are...
This is sort of shared infrastructure, but they're also very different agents that people have built.
They share the same sort of, I guess, kind of patterns around interfaces and the Slack integration on it.
What is the result?
You mentioned developers are the ones that are building the agents, but are there like five agents, 50 agents, 500 agents?
Like what's the rough order of magnitudes?
It's around 30-ish agents built by about the same number of developers.
And the number of users is, I don't have an exact number, but I imagine it's around several hundred because it's not just the developers in the company.
Everybody used them, some even indirectly.
Cool.
So I love the use cases.
And again, they're a very practical example of, in general, building agents.
When people talk about building enterprise automation.
But I love that it's bringing...
bringing it into the heart, into Slack.
But all of this is substantially enabled by both the citation side and the tooling side.
Maybe actually let's talk about the constraints first, the tooling side.
So before we talked about how you kind of whitelist the tools versus blacklisting.
How does that work here?
All of these things, they need different tools.
Do you share the tools between the agents?
Do you define them?
How does that work?
All right.
So the agents all have a shared baseline of tools.
And you can exclude some of them or you can add custom tools to your agent.
Basically, like skills, except they have a, how do you say, a schema that gets validated and you can run code and do deterministic checks.
For example, another example of an agent is an agent that actually doesn't operate in Slack, but it works on Sentry.
So Sentry is an error management platform.
errors show up on Sentry, and we want to, basically same thing as the Slack agents, right?
We want to triage them, figure out the root cause.
If it's a bug, fix it.
If it's an environmental issue.
So these are just basically settings as part of the configuration of the agent, which I guess you can think of as a configuration of the harness, you know, because this is sort of the environment that you're giving them.
Exactly.
It is the harness, yeah.
Yeah, it is the harness kind of with the tools.
Okay, so that's fairly straightforward.
And I guess how often do people, like before you talked about kind of your mechanisms for knowing that you need another tool.
Because that's probably the piece that is, you know, like the dash dash dangerously accept operations, right?
Or sort of the dash dash yellow is also convenient, which is like, I don't know what it's going to need.
You know, can you just let it do its thing and trust it?
So I guess what's the loop here?
You know, what's the interaction model and how often?
Does it happen that people, you know, like it fails because it doesn't have a tool and they come back to it?
Does it typically sort of stabilize?
Like they do that a few times and then generally it kind of works?
Or, you know, what's the experience there?
I think it's actually quite rare already from the, after the first few agents were spun up.
Because we kind of plateaued around 60-ish tools that most agents have.
And that basically.
Because the tools are shared between the different agents.
Yeah, so somebody adds a tool, all the other agents get it automatically, and then it just becomes possible.
So the way that we discover that a new tool is needed, so the obvious choice is somebody would tell me.
That's quite bad because in many cases people wouldn't tell you or wouldn't even know to tell you.
And then there are a few other mechanisms.
So the citations.
They show up in line in the text, whether the interface is on Slack or another platform.
And then what also shows up is the feedback mechanism.
So you get two thumbs up, thumbs down.
And if you thumbs down, you can write why.
So that's the obvious one.
It goes and when you specify feedback, that gets linked up against the original data that backed this claim and not.
the specific answer.
So already the agent is learning just based on if it queries the same data, it's going to see that in the context.
So you have that.
And then if somebody is talking in the conversation after the agent has responded and he's like, man, I just wish it could, you know, give me an update every 15 minutes.
So that gets recorded, right?
It's just on Slack.
And then...
One other mechanism that the platform has is a schedule mechanism.
So the agents can have a trigger every schedule.
And the first one I added as an example really was the dream state, which is where the agent spins up every week and goes and reviews all of its interactions, feedback, and also the responses people wrote just around the answer.
And they might say something like that.
And then the dream state is going to go and submit a pull request to improve the agent.
I like to joke that this part of the agent is the part that replaces me.
As it sort of chooses to sort of evolve it.
Cool.
So this is the loop.
Like, this is the part in which it sort of optimizes it.
And in those contexts, the agent might come back and say, I need a new tool.
Or, of course, it might sort of change its own system.
There's one more thing I forgot.
There's also a tool called submit feedback.
which is for the agent to submit feedback.
So if it's trying to do something, so it's looking at the tool list, right?
Or you ask me for some World Cup results on it, and I don't know where those exist on it.
I need help.
Then that gets saved.
And when the DreamState runs, it's going to see that in the context of everything and possibly propose that change.
Got it.
Okay.
And it's interesting that the tools are shared between the agents.
I guess I expect at some point in the future you might find yourself sort of isolating and kind of further constraining the agents, but just hasn't come up because you're already sufficiently contained.
So there is a bit of divergence between them, but it mostly has to do with functionality because not all of the agents are only Slack bots at this point.
There are other interfaces like the Sentry one I mentioned.
And so sending a Slack message.
is a tool, right?
And going on about control and the agent's being constrained, the agent isn't able to select where it's sending a message to.
It gets triggered in a particular context, and then it can call send message, but it only sends it to the same thread that is unlocked for it.
It can't choose the destination, so it can't make mistakes, provided that the harness only allows it to be triggered in the right context.
But that's deterministic.
We know how to deal with that most of the time.
I think that very much is aligned with kind of my view, right?
I think when you have the agentic stack today, you have the models that are kind of our operating systems we're building on.
And then on top of that, you have tools and context and harness.
And the context is kind of sandwiched, you know, between the sort of the tools and the harness with the harness creating constraints.
And so that drives some determinism and some sort of security.
But then also the tools drive another.
another set of constraints and determinism.
So I like that.
You're sort of building that.
And I love also, as we had conversations about this, it also sounds like you didn't set out to do this, but in practice, this also saves you substantial costs because you're talking about a fairly notable amount of use here.
It's scanning a whole bunch of data.
It's doing all that work.
And I guess have costs stayed contained?
Like it hasn't really expanded?
Pretty much, yeah.
The actual usage is...
quite cheap relative to what I would expect.
And I think it is indeed because it's very structured, right?
The constraints keep it sort of on the rails.
Just setting up the knowledge graph.
which we'll get into, was quite expensive, but it was a one-time cost.
So let's go there.
We've now sort of teased it up enough on it.
So, okay, and this also relates, I think, to the citations infrastructure, right?
So maybe talk a little bit about the knowledge graph, and maybe we can kind of slip into citations.
I know you have that sort of reviewer aspect of it.
So how does that work in sort of the setup?
Yeah.
So the knowledge graph is really...
So, okay, I didn't mention this, but...
Quite early on, when I had the idea of maybe I want citations, maybe I want all my data to be backed by a source reference, I went and read.
I am using, you know...
L quotes here, yeah.
Yeah, I'm using L quotes.
The AI version of read.
I did assisted reading using chat GPT of a bunch of white papers.
And all of these concepts, data provenance, claim verification, they have been explored quite thoroughly.
over the years and especially in the last year.
And it just made sense to put them together in this way, having all of your data backed by a source and then using that to do claim verification and hooking up feedback into that.
It all felt natural.
It didn't come to me initially.
It came gradually.
So the knowledge graph is an extension of that.
I was looking, so the obvious thing, most coding agents or interactive agents, you're going to give them access to data, maybe to an LLM wiki, you know, markdown files, and they're going to go and explore.
Now, as a user, I know that sometimes they will not explore all the things that they should be exploring, and then you're kind of down to, will they find a hint or something that pushes them in the right direction?
And Andre Carpati wrote about this, that having those links in the LLM wiki helps them find the data they're looking for.
But it's really RAG in the sense that using a RAG to build context and just dumping that into the agent write-off is no longer necessary because agents are better at exploring nowadays.
Agenting search versus RAG is...
Yeah.
And I see that, but again, going back to my need for control over agents, I didn't like that sometimes the results would be inconsistent.
And so my way to solve that was, okay, let's take all the data, put it into a knowledge graph, a database really, with interconnections between the entities, almost a rag.
But then don't actually use it like a rag.
Let the agent walk the connections like it would an LLM wiki on disk.
You can have many, many more pieces of data because now it's indexed and the search is fast and finding relations is fast.
This is just to sort of clarify a little bit about what type of data might be in there.
Like in the organizational context.
Right.
So it's anything from our code.
Slack messages, a particular reception popping up, basically anything.
But setting this up has quite a bit of cost because extracting the entities and all that, that goes through LLMs.
And then probably the most important piece was working on data quality, which is still very manual.
Because if I just took all of Sayara's data, the relevant data at least, obviously not all the data, and just stuck it in a knowledge graph, I think there would be mostly noise, I don't know, maybe bots on Slack sending alerts repeatedly, probably dwarfs the amount of messages that humans send, which is the signal.
So figuring out the different data sources that we want agents to have access to, how to model that in a graph.
An account, an exception, a service, those are entities that need to be in the graph.
And then as the data is ingested, seeing where the noise is and actually going and cleaning that up.
So it took me...
Well, manually, maybe AI-assisted.
But yeah, it took me a few weeks to finish ingesting because it was really bottlenecked on my own ability to...
look at data.
I would, of course, look at the noisiest pieces first, like you group by the source and the origin and then clean that up.
But now we have about a million or so pieces of data in the graph as opposed to many, many millions that would have been there otherwise.
And the agents find surprising things, which is a good sign, I think, as to the amount of signal in there.
So you gather data from a bunch of places.
And I imagine once that data is gathered and you kind of cleaned up the sources, those are now repeatable, right?
You're now kind of keeping it up to date with new data as it comes in of that shape, I guess, kind of from the relevant Slack channels and the other.
You gather that up.
And then I guess, what are the tools that the agent uses to consult that?
This is like a data lake, right?
This is now kind of a context lake, a context graph.
You know, you call it a knowledge graph.
You can use whatever words you want.
And you say that the agent is able to kind of mine that data in more efficient ways of doing it.
What tools make that possible?
So there are a few generic tools.
An exact search.
So if it's looking for a particular service, for example, it would use an exact search and find entities that have that.
that value in them.
So that maybe backing up a little bit.
So the knowledge graph has episodes, which is things that happened and got ingested.
And then it has entities and relations between those entities.
So if it's looking, say, at a particular exception or a Slack message, in that Slack message and exception, you have the entities that participate.
So in a system prompt, which is shared to all the agents.
So there's a piece of the system prompt, a block of the system prompt, that's shared to all the agents and explains kind of how to walk the graph and what a serial agent should behave like.
And then you have another piece that's per agent.
So that shared piece explains you should first use find exact, if it makes sense.
If not, use the semantic search.
Once you find a relevant entity, go ahead and walk the graph using find-related.
So that just finds that.
It's kind of like listing the directory and looking at interconnected entities.
So in Carpatty's LLM Wiki, those would be the dimensions within the file.
And then once it does that, it has a bunch of specialized tools per data source, right?
So it might find an episode of a Slack message, and then it has a tool to read that Slack messages.
or look up more messages by that person and so on.
So those are the bulk of the tools, right?
There's probably five or six tools per data source, and then there's the generic tools.
And those are all read-only, and there are a few write tools that involve with actually solving problems, like performing a code review, posting a comment somewhere, submitting a pull request.
Cool.
Good.
So the context graph is a great tool to basically spare a bunch of this work during the agentic run.
And I love that, again, you were driven to create it because you wanted determinism or consistent success when it should have the answer, that path on it.
But it probably is a pretty big time and money saver as well because now it doesn't need to.
you know, kind of blumber its way, you know, sort of through the data, and it doesn't read unnecessary data, or it doesn't repeat that, which I guess kind of comes back to, you could just ask for structured data or sort of cleaning it up.
Yeah, the goal was context engineering so that setting it up so that the right data is in the context.
And I wanted to, for example, if we go back to the very first use case, it was...
finding a piece of data and finding what's related to it.
So if somebody is talking about an exception, I want to see the pull request that introduced it.
I want to see the exception, the logs, the impact it has.
And for the agent to reliably do that, in my experience, it has to have a way to find those relations.
And they're there.
Like if I as a human can go and look at a log line and something like that and find them, then...
All I have to do is build a tool that sets it up.
So when it calls find related, those things just pop up.
Yeah, I mean, I think it's basically the conversation about structured versus unstructured data.
And I think what we're sort of seeing is agents are so powerful that you can give them highly unstructured environments.
You can give them bash and give them files, and they will get a lot done.
on it but when you come to doing things at scale reviewing large volumes of data and you know performing actions repeatedly and you want to improve reliability i i think there there are multiple paths you can take you know there's the path of um you know better instructing the agents to eventually figure it out.
That's like playing.
I mean, I guess you can make significant progress, but you're just playing with kind of make smarter individuals, right, you know, in the elements, but you're not actually sort of equipping them.
There's, you know, another path you have around sort of giving them.
making the data more structured, so making it easier for them to sort of look up the right information at the right place.
And so that is helpful, of course, kind of adopt the context.
And then you can take it further and you can say, okay, now that I have that structured data, I can actually offload some of the cognitive efforts into tools and run that, which makes a lot of sense because it's basically how we operate with humans.
Yeah, I actually really think about LLMs as humans.
I think that...
It's really tempting to think of them as computers, but really the only other thing in the world which we know that if they have the right training and they have the right information, they generally do a pretty good job, are humans, but they're not infallible, and neither are LLMs, right?
So you're really helping them.
Having these constraints and tools, that's kind of how I work with myself as an engineer.
I try to set up my development environment so that the compiler yells at me when I do something wrong.
I try to set myself up so that my own mistakes are caught before they actually have an impact.
So I do that for...
My friend Borg, too.
Yeah.
I think that is a very effective way.
I think about that as well for a bit.
And there's somewhere in between humans and sort of systems.
You know, sometimes you use a system analogy.
And sometimes you use the humans.
So, cool.
So I just want to take stock a little bit, right?
Like at this point, we talk about how you've gone from choosing the citations-based and sort of tools whitelisting approach to create the agent.
a sort of shared repository with infrastructure for everybody to create an agent.
And then you started by hosting yours and giving Slack interactions to it with both the kind of explicit and ambient mode of interaction, which is pretty cool.
Basically, who's in the driver's seat, right?
Like who initiates the conversation.
changes over there.
And then people started kind of evolving and building their own agents.
I really, really like that sort of hack of like, name it and it's yours.
Build those out and expand them.
And then we talked a little bit about the sort of efficiency component, you know, to improve both accuracy.
enable citations and reduce costs in the process of that in the context graph.
So, you know, a ton of evolution on it.
I think we're sort of running out of time here a little bit, but maybe just a couple of things that, you know, when we were talking about this, I found interesting.
One is when you teed this up, like in an organization, you're this engineering leader who's you know, building, being an IC here, building a bunch of things, promoting it out.
But when you permeate it through the organizations, you worked with this AI captains group on it.
So, because oftentimes, again, like affecting cultural change is a recurring need of it.
So tell us a little bit about that group and how does it work?
Right.
So to get started, we created, Tom and I created a small group called the AI captains.
And we basically, picked a few people who we already knew were more, were interested in building agents and stuff like that.
And a little bit before that point, we're talking December or something like that, I built my first agent because I tried to accomplish a task.
It was classifying the amount of time it takes us to fix a bug from the pull request that introduced it to the pull request that fixed it.
And I tried doing it with a coding agent.
It had some success, but it was very inconsistent.
And at the scale that we are getting meaningful statistics out of something so inconsistent was really hard for me.
So I figured, okay, let's try and build an agent and give it hints to keep it on the right path.
And sometimes you would see that.
you would see that the system hints leak into your context.
When you're working with a coding agent, you would have the agent suddenly react to a system hint.
It would say something like, got it, the command response is this way, so don't tell the user.
And it excellently tells you that.
And I saw what a massive difference it made to actually constrain the agent and give it just the right tools it needs to do the job.
And of course, it takes a few iterations, right?
It's not a one-hit wonder.
But the consistency and the quality actually improved.
So I wanted, I felt like the best way to get other people to build agents is to give them that sort of eureka moment as well.
So they would be like, holy shit, I need to be writing more agents.
That's what I felt.
And so we formed that group.
If it was just 10, 12 people initially.
And then we did like an agent workshop.
And it was all really simple.
And then with this agent's platform, they would each spin up their own agent.
And spinning up an agent to get a carbon copy of Borg is really just it's copy paste.
It's a configuration file.
But once you do that, you're really seeing kind of the power of it.
And it lowers the barrier from there's this unknown thing that's building agents to, oh.
There's a Python class I can modify.
Easy enough.
So you share them, you want them to sort of see the light, and then you kind of make it easy for them to get in the game, which I guess comes back to that.
Fork the agent, name the agent, don't change or his agent.
So just lower the barrier of entry and make it easy to solve all the operational problems, which are really not that interesting, whether you run your agent, what are the guardrails like, like having all that figured out, which are...
Hard questions to answer, but not interesting from the application standpoint.
I thought that was like an effective, sort of simple, repeatable, but sort of effective way to promote things.
I think it worked really well.
And itself, it became self-sufficient too.
Like once those first people really adopted it, some of them became massive champions of it within the organization.
And I reinforced that.
So I would tell people, you don't need my code review.
You can ask.
those people.
And there are the guardrails.
They do require my code review, but I added some of them to be the code owners for that as well.
The guardrails are really how we're keeping Claude or other coding agents from really going too crazy and removing some of the safeties.
So I would really reinforce that.
And kind of my approach has been...
All carrots and no stick.
Just make it really easy.
Make it safe.
Really making it safe is making it really easy.
It's another way of doing that.
Avoid concerns about blast radius and sort of concerns.
Exactly.
So people don't have to worry.
And reinforce it when somebody is effective and he takes the reins.
Be like, good.
Yeah, keep going and point people at them.
Sort of celebrate the successes.
And I think the last use case I wanted to just sort of like get from you as we sort of maybe like wrap up a little bit after that is how you kind of dealt with secret data.
So like you're in the world of data security.
A lot of that data is quite sensitive.
And so on one hand, like as a company, you need to deal with the fact that, you know, you have.
You need to secure and sometimes troubleshoot problems when things don't work on data that you're not supposed to be able to see or that maybe there's specific individuals that have been under NDA or whatever it is have the right clearance to be able to see that data.
So you've also used agent a little bit to export your intelligence and sort of importing the data on it.
So maybe tell us a little bit about that use case, which found a good novel use for agents.
Yeah.
That's one of the use cases, I think, that actually we weren't really able to think about before agents were a thing that everybody inside are thinking in terms of...
Before you started building affinity to, like, what is this thing, right?
Yeah, because, okay, I'll talk about it then.
It'll be obvious.
In some cases, you could have, for example, the on-call triage agents, they have access to all the data, and you have access to the same data, and it's easy to see, and it's easy for them to...
and it's easy for you to see that it's correct.
But in other cases, you may be working in an environment where you're not allowed to see the data or even access the data for regulatory reasons or because of customer commitment.
But obviously, there are still bugs.
So you do need some way to deal with that.
And up until we had agents.
Really, the only way to deal with that is to get a particularly small group and pre-approved group of people who can see the data, and they can then go and troubleshoot it.
But as a company like Sierra grows larger and larger, I mean, we grew quite quickly.
I think a year and a half ago, we were massively smaller.
You may need more people than that small group, right?
And how do you do that then?
So we came up with this idea of having an agent that's within the controlled environment and has access to the data.
And you can have it investigate the issue and troubleshoot it.
And it has access to everything, right?
But one big constraint, it cannot output except for structured validated output.
This troubleshooting agent has access to everything, but then it can only output for a fixed vocabulary, which you can verify doesn't leak any data.
And then so it does its investigation, outputs in the fixed vocabulary, and then that can actually go out of that perimeter to other people who are not approved because now you know nothing sensitive was sent out.
And then we have another agent which reads that data.
and has access to the code and all of the wider availability data and can put together an explainer for humans backed on that other agent's conclusions.
So those two agents, they're called Mulder and Scully because they worked on the X-Files.
Hey everyone, hope you're enjoying the episode so far.
Our team is working really hard behind the scenes to bring you the best guests so we can have the most informative conversations about agentic development.
Whether that's talking about the latest tools, the most efficient workflows or defining best practices.
But for whatever reason, many of you have yet to subscribe to the channel.
If you're enjoying the podcast and want us to continue to bring you the very best content, please do us a favor and hit that subscribe button.
It really does make a difference and lets us continue to improve the quality of our guests and build an even better product for you.
All right, back to the episode.
This is basically a way of having kind of two agents talk across a wall, you know, when you sort of can't see the data.
And I guess the result of that is you can troubleshoot issues a lot more efficiently.
Because expertise varies, like the people around it.
You can have somebody on an R&D team that doesn't have access to the data program the agent with all of the knowledge of that team and how to troubleshoot and examples.
And then the agent goes and does that in a controlled environment.
And the feedback is fixed and safe.
And then going back to that escape hatch thinking, the controlled agent can still output.
to the pre-authorized small list of people.
If it does need to say something that it can't say using the vocabulary, it goes and outputs that free text to that group, and then we can go and expand the vocabulary after human review by the approved people.
So it all comes together.
Yeah, you scale those people.
So I thought that was very cool, and I feel not just about confidential data, but also when you think about on-prem support, when you think about all these paths, I think that's very compelling.
In many use cases where you just can't have the data go out, but you still need to operate.
So, Ori, thanks for sharing all this.
I love the journey on it.
I mean, we told it.
I felt like the destination and all those insights are super useful.
But I think the way to achieve the journey to iterate and how did you learn and what did you add later, I think is...
just as educational in a lot of organizations that are going through that process.
So thanks for sharing.
And I guess just before we wrap up, I'll ask you, what are you excited to do next for this whole journey?
So I think right now that the most exciting thing is working on our code harness, coding harness that uses all of that infrastructure.
But honestly, I feel like...
If you ask me two weeks from now, it might be something else because there are a lot of ideas that we couldn't have had before, like this confidential information agent.
I don't think I could have thought of that a few months ago before I had internalized thinking in terms of agents.
So I'm really just excited about that.
It's sort of a meta excitement.
It's not a particular thing I want to build, but I'm excited to be able to execute on them so quickly.
Building an agent once you have all this set up is really easy.
Getting it access to the right data like this confidential agent sounds like possibly a heavy lift with lots of moving parts.
Two days.
Right.
That's nuts.
It's not just LLMs, but having all this set up and also, honestly, that we're a data security company and you know what data you have and where it is and you can let the agents.
access it and know they're not going to go wild.
I think that's really, really important.
And if I can just give one tip to the audience.
Since having AI adoption is really about adoption, it's about making people adopt this new mode of thinking and building agents.
So I would recommend trying to do it with a platform engineering approach.
That's why I went all carrots.
because you're really trying to get people to adopt a new type of agency for themselves, and you can't tell them to do that.
They have to internalize it.
So you have to lower the bar and make it look enticing and make it theirs.
And I would love to say I've thought about all that before, but it's kind of in hindsight.
I think it makes a lot of sense.
And I think we've learned a lot from the DevOps transformation and the value of the paved roads.
And you enable that.
And I think the other advantage I'll sort of highlight is that it...
helps people along on the journey because you've paved the road and it's easy for them, but it also still allows them to sort of bushwhack their way kind of along the side and maybe find shortcuts.
And so most people will just opt for the paved road.
And then the few that goes to the side, some will regret and come back.
And then some will find a new path.
And so there's an opportunity for learning.
Yeah, you can learn from that.
If you just sort of centralize and you constrain as opposed to like you just work with the sticks, then the result of that is you're saying, well, it's only these sort of smart.
people in the platform team, they're the only ones who can innovate, and the rest just sort of apply it.
And so enabling a paved road, but allowing someone to go to the side is a powerful path.
I think it's easier to have that mode of thinking as the platform team for this, because agents are new to everybody.
So it didn't require any mental gymnastics on my part to assume that there are things I didn't think about.
Cool.
Well, Ori, thanks a lot for coming on the show.
Love what you've been building.
Looking forward to explore the futures.
And thanks, everybody, for tuning in.
And I hope you join us for the next one.
forward slash community to learn more and I hope to see you there.
