# Optimizing AI Coding Agents for Secure Development

**Podcast:** The AI Native Dev - from Copilot today to AI Native Software Development tomorrow
**Published:** 2026-02-25

## Transcript

Too much context is always a problem.
Most of my skill.md's are built by my agent, which means they're overly verbose, and I get way better performance when I kind of cross-reference things.
So whenever I think I have some large concept that is important, I'll say, put that in a new file, reference that in the skill.md, and that improves a lot.
There's three pieces here.
You've got the tooling, agents, and then you've got the LLM under the covers.
And I think people probably...
think too much about the model sometimes.
I'd always argue with people about models and I'd be like, this model's better.
And I would argue with people about that, but why?
The Reddit flames, the flame wars on Reddit of this model's better and that was not better for me.
Even the trust me bro benchmark.
Yeah.
I don't trust them.
I like that.
Trust me bro.
Tell us a little bit about what CodeGuard is.
CodeGuard is really, you can think of it as security skills for your AI coding agent, right?
We use...
All kinds of tools in Cisco.
It's how do we ship that into WinSafe or into Cursor or into Cloud Code with low friction, right?
That's kind of the idea behind CodeGuard.
That's what we're trying to love.
The LLM or the agent will run those scenarios in two manners.
One with the skill, one without the skill, just the baseline, we call it.
CodeGuard did really well.
You can actually see, I think it was 1.79 times improvement on the baseline.
Were you shocked, surprised at that kind of number?
I definitely was.
I was skeptic on, do I really need this?
Well, we...
want to give our agents all these tools.
Are they secure?
Are they exfiltrating data?
Stuff can be really hard to observe when you have an agent that's just going crazy, right?
It's all beyond just traditional software.
What advice would you give to our listeners who, if they wanted to create a skill from scratch today, how would you go about it if you were doing that from scratch today?
It's really...
Before we jump into this episode, I wanted to let you know that this podcast is for developers building with AI at the core.
So whether that's exploring the latest tools, the workflows, or the best practices, this podcast's for you.
A really quick ask, 90% of people who are listening to this haven't yet subscribed.
So if this content has helped you build smarter, hit that subscribe button and maybe a like.
All right, back to the episode.
Hello and welcome to the AI Native Dev.
My name is Simon Maple and I'm at Cisco Live today in Amsterdam and joining me is John Gretzinger.
How are you doing, John?
Great, Simon.
How are you?
I'm doing very well, thank you.
And today in this episode, we're going to be asking the big question.
Can we guide coding agents to write secure code out the box?
And can we encourage and teach coding agents to be able to spot vulnerabilities from existing code bases or code changes?
We're at Cisco Live today in Amsterdam and we...
Just at a session just earlier on.
Your first Cisco Live, your speaker.
First Cisco Live.
It's really empowering.
Finally, finally a Cisco Live session.
Tell us a little bit about what the session was today.
Yeah, I mean, the session was kind of my history as a developer trying to wrangle this world of AI coding, right?
In the enterprise space, particularly, because it's challenging to break secure code.
So it's just kind of walking people through my history and the last two years to how these tools and models have developed.
Awesome.
Awesome.
And tell us a little bit about who John is.
You're a principal engineer at Cisco.
You're part of the CX engineering team.
Tell us a little bit about what you do day to day.
Yeah.
My day to day is a lot of AI focus, of course, but it's, you know, kind of a newer team, you know, like two years old trying to figure out how we bring more value to our customers with AI, right?
And it's building agents for our customers, building platforms for our customers, just really trying to enable the Cisco technology because there's a lot of it.
Technology is complicated.
Can AI help us and help them?
That's kind of what we're trying to do.
Awesome.
So we mentioned CodeGuard at the session today, and I think there's a couple more sessions tomorrow and the day after with Omar Santos.
And of course, Omar, obviously one of the proud owners of CodeGuard.
Tell us a little bit about what CodeGuard is really there for and the journey it's been on.
Yeah.
So I mean, CodeGuard is really, you can think of it as...
security skills for your AI coding agent, right?
So, you know, we really want our developers in Cisco across all of our orgs to use AI coding because it accelerates software development, right?
We want to make more software, but doing that securely and to our standards is very important, right?
And so our security and trust org is trying to wrangle all that and figure out how do we enable our developers without shooting ourselves in the foot?
And so CodeGuard was kind of the first kind of attempt at that.
Internally, you know, it's just kind of a list of skills.
What is secure development?
What are the kind of things we need to look for while we're writing code?
And then the real problem is how do we ship that to all our developers, right?
We use all kinds of tools in Cisco.
It's a company of that position.
So we don't all use the same tool.
Love to, but it's just not the reality in Menden, right?
So, you know, how do we ship that into WinSafe or into Cursor or into Cloud Code with low friction, right?
Now, we don't need to add more friction for the engineers.
So how do we just make that easy?
That's kind of the idea behind CodeGuard.
That's what we're trying to love.
Awesome.
Awesome.
And I think it's such an important problem because people who use agents, they're stung by a number of things.
And one of them, of course, is secure code.
And the biggest problem, of course, is agents have learned, LLMs, they've learned on all of our bad habits, our worst practices, and they average them out and we tend to get that back.
So it's really, really important to provide that guidance.
sharing that knowledge around that.
Let's talk a little bit about, we mentioned so many developers at Cisco, obviously a huge organization, huge enterprise.
How is a company like Cisco embracing AI today across its development organization?
Yeah.
I mean, I certainly can't speak for our whole company.
It's very massive as you meant, but from what I'm seeing, I do internal training.
So I get some visibility cross org on what people are doing, but there's pockets that I don't.
To summarize, I mean, our executive leadership really enables this.
Like they want us to embrace this and learn it.
And we know it's a new technology, so it's not all figured out, right?
There's no roadmap to follow for this.
So they let us get creative and try new things, but obviously we have to do that securely and without blowing the budgets, right?
So I mean, it's a fine line to kind of balance.
But I feel very empowered as a developer in Cisco to try these new tools and figure out productive patterns, share that back with my organization or cross-organization and stuff.
It's kind of the culture we have kind of bring here.
Yeah.
Yeah.
And I guess, you know, in the session that we had earlier, we kind of talked about, I think you talked about the Stone Age and going forward and like the dawn of civilization where we are today.
And it's really like, you know, this progression of learning.
What would you say are the kind of the things which almost like unlocked in your development environment a new way of working?
Was it tools?
Was it practices?
Was it...
you know, methodologies?
What would you say are the things?
Yeah, definitely a combination.
But I think the tooling is key because early on, you know, it was just a chat window, right?
And there was only a few models and people weren't really using them for code quite yet.
Like I was very much heavily into vibe coding back then and copying and pasting between, but it was so clunky.
And I was like, there's got to be a better way to share this.
Like I want this engineer to try this.
And that stuff took a long time to figure out.
And I think.
Really with the agentic IDEs is really when I started to see that.
Oh, this is what I want.
It's like this is magic.
You know, I make that plan that specification.
I'm always about that now.
Curses, windsurf.
Yeah, yeah, of course just specifications even your guys tile is yeah, and now it doesn't matter what yeah, what idea is that that tile helps a lot But yeah, that's just the tooling was move.
I think probably the biggest break to the models obviously That's and that's kind of a given that we knew the models would kind of get better at this stuff but I think the tooling much better.
Like if we take the tooling we have today and we used it two years ago, it would probably still be really good, right?
Even with the older models.
So like for me, that was the difficult part to navigate.
It's like, how do I do this effectively?
It felt like I was getting burned a lot at the fire.
I was like, oh, shiny new tool.
Yeah, it's an interesting discussion actually because of course there's three pieces here.
You've got the tooling, you've got the agents and then you've got the LLM under the covers.
And of course, yes, there is a...
level of LLMs within the agents as well.
But it's interesting how sometimes we just think of it as one experience, which I suppose it is, is that one workflow.
But really, there are a number of aspects here that all contribute to a successful development workflow.
And I think people probably...
think too much about the model sometimes.
And I think you're right.
If, you know, with a good user experience for a developer with the IDE or with the terminal even, and a good agent that actually walks you through, that does the planning, that does, you know, all these different types of things, those kinds of things, I don't know, don't know about you, but for me, are almost more important than the model in some cases because it allows you to perform that workflow.
Otherwise it's more of a bot if there's no workflow.
Yeah.
It felt that way to me as well.
Yeah.
Like I'd always argue with people about models and I'd be like, this model's better.
And I would argue with people about that, but why?
Show me why.
The Reddit flames, the flame wars of Reddit of this model's better and that was not better for me and blah, blah, blah.
Even the trust me bro benchmark.
Yeah.
No, I don't trust them.
I like that, the trust me bro.
That's the only type of benchmark I'm listening to now.
It's the trust me bro.
Okay, so let's go into a little bit more about CodeGuard now.
And we'll talk about, I guess...
How that evolved when you started, it was, I guess, you had a ton of best practices, a ton of instructions or guidance.
How did you go about turning that into what it is today?
Yeah.
So, I mean, the security and trust org, that's the code guard kind of evolved internally.
And I think they took a lot of like OWASP best practices and industry standard stuff that we already apply in Cisco internally, but just kind of reducing the context.
Because if you've ever tried to go look at OWASP rules, it's very encompassing, right?
And it's not.
something you can even feed to an agent that would make sense of that to make sense to your code so it's just simplifying it right even for humans for agents we need the same simplification so and then packaging it right so they wanted it to work everywhere it's just text right i mean at the end of the day this is context we're shipping they um you know internally we work on different methods to do that i think just a repository of code is is a simple way that everyone can understand but then it's there's philosophical debates on do i commit this to my repo but then what if it changes right Wrangling that world is still not solved.
It's interesting what you say there about the fact that these are similar, they're basically OWASP rules, right?
Nothing's changed in terms of the security space of what we're actually looking for in code.
Is there anything agent-specific in there, or is it entirely just about the, could you run this against human-written code as well?
Yeah, of course.
I mean, it's mainly for just traditional code, right?
Then agents are creating traditional code, but they do also, I think it was two weeks ago or so, they added MCP security, right?
So that's obviously top of mind as well.
We want to give our agents all these tools.
Are they secure, right?
Are they exfiltrating data?
Yeah.
That kind of stuff can be really hard to observe when you have an agent that's just going crazy, right?
There's also MCP security built in as well.
And I believe skill security is coming as well, not into that.
So it's evolved beyond just traditional software at this point.
And for the distribution of this now, you obviously want all of your users to consume the rules, the guidance.
And the way that you did that was you built it into skills, which you made custom for each of the different agents.
So there's a dot cursor rules, there's a clawed skills and things like that.
What were the kind of problems, I guess, that that caused when you were trying to satisfy so many different agents and their different structures?
Yeah, I mean, I can really speak as me as a user of CodeGuard, and I actually use all those tooling.
So I speak from the experience of trying to install in each one.
I mean, it's really just kind of a package.
You would unzip it, and if it's for Windsurf, it unzips in that location.
I mean, that's fine, but again, maintaining that's very difficult because I wish we would just standardize on this stuff in terms of agents.md, right?
But even some IDs that claim they honor it, they don't, and that's really the problem.
And maybe Windsurf would start.
to solve that standard and then they don't do their old rules so now it's maintaining all that context every developer having to do that just doesn't make sense like it's just too much friction people just don't do it and it just makes me not want to use code guard because it's too hard to keep up to date right and that's not you know that's bad well i want it to be easy just always be up to date and just easy for my agent to to pull in you know that's what i want so let's talk about learnings then so obviously you created this what were the things that you took away from it straight away from uh from using it and playing with the structure of the code guard skill.
Yeah, I mean, and it was so specific to the tool that you're putting in, of course, but, you know, so even I had to go, I learned stuff on how they deployed it.
I'm like, oh, I was doing this wrong when I was using this IDE, right?
So, like, that stuff was great to, like, it's a learning experience, but at the end of the day, it's really just a bunch of, like, markdown files that are intuitively titled for the security practice that you're trying to do.
Not every security practice applies to every piece of code or repository, so.
Being able to kind of pick and choose, oh, I'm trying to do this.
Yes, session management's a big problem I need in this repel or SQL injection, right?
Just take those two skill files.
I can read it as a human to understand what it's doing, but just feed it to the agent and have it review my code.
And they can have a conversation about it with my agent because I don't always agree with my agent, of course.
And sometimes it makes assumptions that are wrong and I have to steer in a direction.
But with security, I want that to be easy.
I don't.
No one knows everything about the different ways you might exfiltrate code or exploit code.
And what about the repositories that people who were trying to pull CodeGuard into their agents, how did it affect other people's repositories?
Yeah, I mean, there's a lot of opinions there.
And that's kind of the problem.
Some people say, this should be committed to your repository.
So it's always there.
But if you have a team of developers that use all different tools, you're just bloating your repository with all these .files for each IDE.
Even I myself wrote a script.
Claude wrote a script.
That is like an agent's linker.
So I would write just one agent's MD, and then I would run the script, and it would create all those directories and simlink it off back to my single search of truth.
But it still bloats my repository with a mess.
And so I don't agree with the philosophy that stuff like Kogart should be committed to your repo because...
every time it updates, you have to merge it in.
It really complicates PR processes too.
It's like just noise that your team shouldn't have to deal with because it's not part of the central code base.
I feel like a repository should really reflect what the repository is about and not external stuff.
Just like NPM packages.
I don't have the whole NPM package committed, right?
That would be crazy, right?
So same concept.
I really think it should be treated that way as you guys do.
Think of it as HESL.
Yeah, absolutely.
And now when we think about the skill being used by the agent, identifying issues in existing code or trying to find issues or trying to write secure code with the guidelines as the best practices.
What were your immediate learnings from how it was working and how did you know it was as good as it could be?
Well, I mean, I didn't.
I think it was a lot of why did you use the skill this time?
Why didn't you use it this other time when I wanted it?
Sometimes I'd have to go out of my way to be like, but did you even use these skills and CodeGuard for this?
And it'd be like, oh, no, I didn't do that.
Let me go do that.
And then it would change the code.
It's like, oh, he's wasted my time.
So there's definitely a lot of that with any skill development, not even just with CodeGuard.
I go through that a lot.
Typical activation of skills and with any agent.
It's not always easy.
Of course, you can try to set up hooks and stuff in Cloud Code and make it hook off that kind of stuff.
But even that's not perfect and it can be kind of...
it can actually be worse.
Do you have any tips or best practices about how you can increase the activation?
Is it something that's in the skill as a producer of the skill that it's good to do?
Or is it something more as a user, you want to guide, you want to tell the agent, I need you to do these things in this order?
What's the...
Yeah, I mean, I use, when I build the skill, the things that are top of mind for me from what I've just personally learned is keep the skill lean.
I think it's a best practice anyway from the topic.
But that's something I've learned.
Like you can't bloat your skill.md.
it's just going to become useless.
So I think referencing things in the skill.md is really the way to go.
And when the model doesn't pick up on key things, like when I'm having a conversation, actually, I just ask it.
I just say, why didn't you pick up this hint in the skill.md about the security review, right?
And then it'll explain to me, and I'll just ask it, can you go update your agents.md or the skill.md in a way that next time you go to use this, you'll actually pick it up when you do it.
So I tell the model to fix itself because- It's like self-healing kind of flow, yeah.
Who better to ask than the model itself?
So that's my strategy and it works quite well.
And also it's less thinking for me.
So it's just like, why'd you do this?
Oh, okay.
Well, don't do it that way and change it for yourself.
Yeah, yeah.
It's a lazy approach, but it works surprisingly well.
And how much can you, you mentioned about putting things into the Agents MD and things like that with context like this.
Obviously there's a huge amount of context and it's quite easy to blow the context window pretty quickly.
Did you- Did you play much with the context window of how much you should put in, how much is mandatory for it to read versus just allowing it to be reference and you can play with this whenever you feel as an agent you need more information on this?
Was there a balance there?
I don't know what the balance is.
I'm still trying to find it myself.
We're all learning.
I've certainly tried, but I don't know that I have anything other than anecdotal memories.
And I don't have a lot of time to spend engineering context management for my age as much as I want to.
I got to write actual code, right?
I ship software, not play with AI coding and, you know, babysit it, right?
So it's really a trial by error learning thing.
I don't have time to really engineer it, which is unfortunate because I think that's really what's needed to get the optimal performance out of these is really engineering the context and how it's picked up, how different models key into that, how the different tooling actually is plugged in.
It's a very complex, you know, matricity from that.
And I just don't have...
the time to really dig into that as much as I would like.
Yeah, yeah.
So what were your learnings then once you had the CodeGuard skill?
What were the lessons that you kind of went through in terms of getting that skill as good as it can be?
Yeah, I mean, learning just the basics of model activation and that stuff.
I mean, not even just CodeGuard, but I've written a lot of other skills myself, which are...
I spend a lot of time trying to understand what works well and what doesn't.
Too much context is always a problem.
You don't want to bloat that skill.md with just unnecessary information.
I mean, most of my skill.mds are built by my agent, which means they're overly verbose.
So I found that problem very early.
I think it makes way more sense, and I get way better performance when I cross-reference things.
So whenever I think I have some large concept that is important, I'll say, put that in a new file, reference that in the skill.md, and that improves a lot.
But it doesn't always pick up from the skilled at MD to reference that when I should.
So it's still learning how the best way to kind of do that in each model is very different from that perspective.
So it's not been easy.
And in terms of how you know the context is right, before you ran evals, was it anecdotal?
Was it gut feel that this is correct?
Largely vibe feel.
Oh, vibe feeling.
Did I get a good vibe that this worked well?
Did it do the tasks that I wanted to do efficiently?
Or did it stumble around forever and 10 minutes thinking or whatever?
So no, it's very anecdotal.
I still am trying to figure out the best way to really evaluate these more like software is evaluated.
It's not easy to do.
So yeah, I don't have a good data backed way to...
So what did you, was there anything you tried?
Did you try using LLMs to judge it or anything before the Tesla evals?
So I'd say I shared it with other engineers.
So, you know, only when it was like, oh, I did it two times myself and I got a good result out of it, then maybe I would share it with someone else and get their anecdotal feedback.
But it's anecdotal, anecdotal.
Anecdotal at scale.
Yeah.
And then, of course, they would have, they're like, oh, I use a different tool and a different model.
Right.
And then they'd have way different experience.
I don't have time to go.
Which could actually be nothing to do with the skill.
It could actually be the agent, the model, and a number of other things.
So we ran with Tesla evals.
We ran with the review eval and then with the task eval.
Tell us a little bit about the journey you went on with the review evals first.
Yeah, yeah.
It was actually fun.
It was a good experience, honestly.
I didn't expect you just ran Tesla skill review.
instantly within just like a couple seconds.
I don't think it was very long, but I always assume agents will look at myself.
I might've walked away, but it was very quick.
I came back and I have this rubric to review, right?
Not just categorized feedback, but more importantly to me, actionable feedback, right?
These models, when you ask them for feedback on things, they tend to just judge you harshly and they never provide suggestions.
I mean, sometimes they do, but they give you too many options.
But this is like grounded in, oh, well, this part here is wrong for these reasons, or you might be able to improve for certain things, right?
I might not agree with those categories, but for me, it's an iterative loop always, right?
I want the model to do as much as possible, but I still need to be in the loop because I can't just review, let it review itself and keep going because that's a waste of money and I'll probably end up getting something worse in the end.
So I make myself part of the review process and I'll just kind of look at that rubric and say, well, I don't agree with this point.
But I really agree with this point.
So let's focus on that one right now.
Can you suggest some changes to improve that part of the rubric?
And then it would adjust it and re-eval and then they get a better score.
And it's like, oh, OK, we can start to see the improvement working lives.
And you understand it as well because you're in the loop.
So that's kind of how I do it.
I really have my model improve its own skills because at the end of the day, it's the one using it.
That's awesome.
So you're essentially managing, you're overseeing the updates.
You're recognizing.
having the time to agree if you want to make that change but then the model does it the model and then of course tassel i.e the models in the background they they retest uh provide the new feedback and then you just loop through that until you get to a stage where you're you're comfortable yeah yeah exactly how did that differ from i guess the task evals uh the task evals for those who didn't see didn't hear the last couple of episodes the task evals are much much more real user case scenarios.
So a number of scenarios that are based on what the skill is likely to be used against.
And then the LLM or the agent will run those scenarios in two manners.
One with the skill, one without the skill, just the baseline, we call it.
And then you get different percentages.
Hopefully with the skill is a higher percentage success rate based on the evaluation criteria.
CodeGuard did really well in that.
You can actually see, I think it was 1.79 times improvement on the baseline.
So were you shocked, surprised at that kind of number?
I definitely was, for sure.
Because I think anything with AI, someone's like, use my context, it'll improve your skill.
It's like, none of that comes with data other than the trust neighbor benchmarks, of course.
But I think I was skeptic on, do I really need this?
Right?
Like, I know security.
I can just tell it.
you know, go read about security and do this.
Like, where does this help me?
And it is optimized, right?
So it's way faster.
And the 1.8 improvement was impressive, especially when you have a baseline to compare it to.
Yes.
Then you know, like, oh, well, what is it actually better at?
I don't believe it.
It's like, oh, I see that.
You can go try it yourself and you see that.
And the joy is actually as well.
This is kind of like very early stage with the task evals, but the scenarios are there to be reviewed and updated and then sent back and re-evaluated.
So if the evals aren't accurate or aren't realistic, It's absolutely changeable so that the skills can be tested against valid scenarios.
The baseline was 47%.
So 47% is the score from the evaluation run criteria that a plain clawed code did against the scenarios.
agent success, so with the skill.
In some cases, I'm looking at scenario five here, where without context, everything is red.
Everything is X's on the left-hand side.
With the context, four out of the five turned into 100% success.
One, still 0%.
What do you do with this data?
How would you turn this into actual valid turnaround for your improvements?
Yeah.
I mean, it's a good...
I like the scenario-based, especially because...
Not everyone needs the same security, especially like the session fixation one was especially interesting.
I think that's the scenario five.
One, because using the same session ID after you log in is sometimes a bad practice.
How long does that ID live?
There's different schools of thought on that.
I mean, us in Cisco, we are very zero trusting.
So it's like that's very short lived.
You need a new one at least once a day.
But, you know, for maybe more consumer side stuff, that's annoying to log in so frequently.
So maybe you don't care about that.
So having a different scenario for enterprise versus like more consumer base and being able to tell like, is my agent aware of that?
Is the base model aware of that?
And if not, how can I have this skill kind of give it that context?
Because again, perfect security doesn't exist, but, you know, there's different levels, you know, security will slow you down.
So you don't want it to slow you down, but you want it to be secure enough for your application.
Yeah, absolutely.
What would you say were the biggest learnings from the point of view of how useful you found the eval data in certain cases and what learnings you then had from CodeGuard itself?
Yeah, I mean, it's definitely even some of the scenarios that it helped create helped me think, oh, there are situations that I didn't think about.
that it kind of applies.
So it really helped kind of, I don't want to come up with all the scenarios.
I might not be thinking of them.
So from that perspective, it was good to see, but really the rubric of how it thought about it and the way that it explained it is like very different than how I would approach it, but it's a better way, right?
It's more scientific, which I like.
And I feel sometimes I'm just exploring and I'm like, I think this helps and I don't know.
And I really hate that feeling that I'm wasting my time and I'm not even making improvements.
It's the worst to me.
So I think it really helped me keep on track.
give me a hill to climb, right?
Like that's what I try to do for all my evaluations for my own agents is what are we trying to build?
What does good actually look like?
I think this gives you a example of what good does look like or at least an explanation of what good should look like and then where you're missing the mark.
It makes it easy to just kind of fill in the blank, climb that hill and get to where you need to go.
It just makes it less friction.
Yeah.
Way less friction to get there.
Yeah, yeah.
And using something like this with a skill...
Would you say, you know, for a Greenfield project, writing code from scratch, the skill was maybe better for those types of scenarios?
Would you think Brownfield, where you have an existing set of code whereby you actually want to do more code review or slight adjustments to the code, where would you say a skill like this actually has its most effect?
I mean, everywhere.
That's the joy of this type of skill, actually.
It is, but I would say...
I don't want it to get in the way too much for simple projects, right?
There's some side projects I want to do and I just want to code and I don't care about security, right?
If it's for me, I'm the only one using it and it doesn't matter what the data is, sure, I'm just having fun.
Then I don't need this, you know, but then it's like, oh, I made something fun.
I want to share this.
Now I got to go add this security stuff to this, right?
So, you know, when I'm just having fun, maybe I don't use it.
But for anything I'm actually going to share or ship, I think it absolutely has its place always.
Amazing.
So what advice would you give to our listeners who, if they wanted to create a skill from scratch today, how would you go about it if you were doing that from scratch today?
Yeah, I mean, I know that I have the best approach, but how I've, where I am today, probably different in three months, but it's really building it with the model together.
I don't, you know, write the skill.md myself.
Obviously, I don't have time for that, but also asking the model what it thinks, right?
in terms of how do you think we should do this and i build it together right not only does it help me understand some of the stuff i didn't but it helps the model understand what i'm doing because oftentimes i'll say we're going to go do this task right and it makes a lot of assumptions on what's involved in that process and that's wrong so oftentimes i don't make the skill first i'll basically tell it i want to make a skill and then i'll say you're going to do a dry run with me we're going to do this together and then at the end you're going to review everything we just did and we're going to make a skill out of that And then we also have our first evaluation from that as well.
It's a real world scenario, me using it to do what I want it to do.
And so then I tell the model to build this, go off with that, and then I adjust and improve from there.
So that's kind of my approach right now.
It's very hands off.
I really love that approach.
So you're essentially having the LLM while doing the stuff.
Observe the interactions, I guess, because if the agent is creating something and you actually say, do you know what?
That's not how I want it done.
I want it done like this because of...
blah then all of a sudden it learned ah okay my my my essentially my basic route was to do this you didn't like that as a result i'm going to start adding this as part of the skill absolutely love that um and also if it's doing things the right way and you actually don't comment on it it actually doesn't need to document it so long as it is deterministic enough to do something similar-ish every single time right i i love that um and then and then evals and things like that how would you go about uh running evals, when would you run evals, I guess, on that kind of a skill?
How many iterations would you do before you'd be happy with the skill that you'd want to eval?
Definitely still figuring that one out.
But I mean, I'm a big proponent of evals.
So I say eval always.
But obviously, don't have budget to eval on every line change.
So it kind of depends, right?
But I think major refactors of the workflow where I'm changing something in one scenario that could affect another and I'm not sure, you know, I run it.
But again, it's...
really use case specific.
I don't have a really general guidance of when I do that, but I'd say it's kind of like if you're doing semantic versioning, right?
It'd be like the middle version release, right?
You want to do it at the third release, you know, but again, it's cost and performance balance.
And how about, you mentioned obviously, you know, a large organization such as Cisco, you're going to have different developers that are potentially...
working in slightly different ways but similar policies, similar overall team ways of working but they're potentially using different models and different agents as well.
At what stage if you were building a skill, let's say for a larger team, would you do it a similar way or at what stage would you try and bring external thoughts in because you don't want it to necessarily be a perfect skill for you but then actually...
Not great for a poor nut.
Yeah.
I mean, again, it kind of depends.
I like to bring the users that are consumers in as early as possible, but that can also be too many opinions too early, too many chefs kind of problem, right?
And sometimes it works if someone just hacks it out for a bit, finds what works and what doesn't, and then kind of takes the first go at it and shares that with others.
And then you kind of build from there.
But it needs to be easy to explain how you got to where you are.
And if people don't agree, they can try something else.
But too early might be not good to have a lot of people, I'd say.
It could be a wasteful of time.
So potentially two loops.
One where it's like you learning with the evals, kind of like refining, refining, but there's no point in refining it too, too much before you get that more external loop of feedback with the wider team or even, I guess, you know, extra agents or extra LLMs to get that extra.
Yeah.
Yeah.
That's one approach.
The other approach is just, you know, have five people or five models all do it asynchronously.
And then we merge the results, right?
I mean, if I'm in a hurry, maybe that.
But again, it depends on the skill.
Awesome.
So what's next for CodeGuard then?
Yeah, so we actually donated it to the Coalition for Secure AI, right?
So it's kind of like a Linux foundation, but for AI coding security stuff.
So yeah, we want to share it with the world.
We found a lot of zero days with it.
It's obviously valuable.
and we just want that to be easy for everyone else to use.
And then of course it's on the TESOL registry as well.
We had a little play with it just to...
just to be able to run the task evals and the review evals.
So right now it's at cisco.com forward slash software-security.
It may not always live there, but if people wanted to take a look at it and see what we did as a first tile, it's all there for people to have a play with.
And of course, if we move it about or check the name changes, et cetera, we'll let you know.
But yeah, that was...
Yeah, people should give it a try.
Pull it down and see if it can find a vulnerability in your repository, right?
I mean, that's...
Why not?
Give it a go.
Awesome.
John, it's been an absolute pleasure and let's go and continue enjoying Cisco Live.
But thank you very much for the session that, well, I say we gave, but you gave like 90% of it, but that we gave.
You kicked the ball into the hole.
I was heading it in on the line.
Thank you very much for the session.
Real pleasure to be here with you.
And thanks for the session.
Yeah, thanks for having me.
Thanks everyone for listening and tune in next time.
