# AI-Driven Energy Data Engineering and Asset Optimization

**Podcast:** Engineering with AI
**Published:** 2026-09-14

## Transcript

what we're seeing is that there's really a brand new role that's coming out, which is an energy data engineer or an energy data software developer.
Their specific function is around how do I manage all of this energy data and build tooling that allows for better organizational decisions for managing these renewable assets, right?
The asset performance space is becoming a really big thing among companies that care about the financial performance of their renewable assets.
So this particular system tool is for those developers that wants to be able to use generative AI to very quickly build energy data apps.
So this actually is kind of a callback to what I was doing in healthcare when I was at Cerner, where Cerner had its base platform, but then they were going into ICUs or they're going into specific units and building on top of that platform using internal APIs, custom UIs and custom workflows that are idiosyncratic to that particular user base.
So this is kind of the same thing, where we create this tool or wrapper that allows people to build custom tooling using Gen.AI on top of their existing data stacks and data platforms.
So they're not having to do all the data plumbing.
Chingi Samudzi, previously AI product development at Google, head of commercial solar strategy at Lumen Energy, has got extensive experience at the intersection of policy, data, science, and go-to-market strategy.
He's a Carnegie Mellon graduate with a background in decision science and international relations.
Currently, he is the CEO of Asoba Systems, a company in the tech space with deep knowledge and everything from solar economics to AI.
I met him when I was at Odyssey.
We were both working in solar and Africa and got to know each other then.
Asoba has been doing really cool stuff in ML for a while, but when he showed me his latest work recently, it blew me away and I had to have him on the show.
So today, Shingy is our guest on the Engineering with AI podcast.
Shingy, thank you so much for coming on the show today.
Hey, Kyle.
Always great to talk to you.
Yeah.
So we just heard your bio and I hinted that you've been working on some pretty cool stuff.
But what would you add?
Like, I think I called Asoba a tech company working in the intersection of energy and AI, but that's probably broad.
Like, what would you say about Asoba?
I would say that we are an economic development company.
Nice.
You know, as much as we get put in the box of climate tech and, you know, I'll cop to that charge as well.
But I think.
For me, the main benefit that we're seeking is really around economic development in Southern Africa.
Particularly that's part of the world because that's where my family is from.
I still have a lot of really deep ties to that part of the world.
Yeah, yeah, right.
No, I mean, it's funny.
When I learned about the energy sector from the inside, I realized how much development is a play.
And I didn't mean software development.
It's almost more like real estate development.
You know, you're going to buy a patch of land and you're going to improve it.
by putting stuff on it.
And, you know, you're going to hopefully create economic value.
So, yeah, that's a perfect description.
I like that.
Are you guys actually, like, building energy plants then?
Or, like, tell me how that part's gone.
Yeah, I mean, so we dipped our toe in a little bit of that entirely in the United States, where we actually were able to get some funds from the USDA in partnership with another company called the Black Table in the Atlanta area.
And we had a thesis around something called agrivoltaics, which is essentially the combination of solar within agricultural purposes.
And there's some really interesting, surprising support from the current administration in the U.S.
for solar in agricultural context.
But there's a whole political game behind that that we can maybe get into.
Probably not as interesting, though, as the stuff that we're doing in Africa.
Got it.
I remember ML being part of your story like way back.
But yeah, like I say, like recently it seems to have really taken on a new shape.
Do you want to talk a little bit about like how, where you came from and what you're doing now and how you built on that?
Yeah, I mean, I would say that the journey into machine learning for me really started at Carnegie Mellon.
So I graduated in decision science and it was a relatively, I think it was a relatively new major at the time that I was there.
So I think I selected that major in, 2005 or 2006, I saw a random bulletin in the department about the specific major.
It really aligned with my disillusion with classical economics at the time because I'd gone in with the intent of really going into developmental economics, which from a Carnegie Mellon perspective, they're much more focused on industrial management.
So it was kind of misaligned, but it was a beautiful, serendipitous mistake on my part.
Because what that led me into was seeing classical economics, seeing the supply-demand curves and then matching it up with oil prices in West Africa and thinking like, well, what the hell?
If this can't explain this phenomenon, what's the whole point of even saying supply-demand dirt?
Right, right.
That is a good point.
Yeah.
But it doesn't really match.
So we had to understand like what are the real...
mechanisms behind actual decision making.
And that was the beautiful part of Herbert Simon, who founded the decision science, probably a field of study, but certainly the department at Carnegie Mellon after he won his Nobel Prize.
So really getting into behavioral economics, rational actor theory, where the elements of irrationality come in.
cognitive bias, how that makes an impact on how people make decisions, and how does that aggregate into the organizational level?
So how do organizations, when you look at whether it's a country or whether it's a business and you say this business is doing this or this country is doing this, A, understanding that it's actually a combination of a bunch of different competing strategies bounded within a specific network boundary that force aggregate decisions in one way versus another.
And when you start to layer that onto the economic phenomena that you see, particularly in my interest area, why there's persistent underdevelopment in certain parts of the world, it started to become more and more clear that this actually wasn't an economic development issue, that it was really a political issue.
People were essentially just choosing.
I mean, I'll be very pointed about South Africa, right?
People freak, like the current state of South Africa where 13% of the population are net taxpayers, right?
That is a political decision.
that people are making to not invest in specific communities, to have them fight over crumbs in these pathways that have progressive ideology labeled all over it.
But in reality, it's really just like crumbs that people are fighting over.
And that's what drives the xenophobia, which actually benefits the people that don't want to share, you know, anyway.
So it really just got, maybe help me understand though, instead of like a...
why is all there's, there's all this inefficiency, like what's going on?
We can make better systems and really thinking about this as like a broken system problem to really understanding that this is, this is actually a psychology problem.
This is actually like in the, in the words of Robert Nozick, a utility monster problem where somebody feels that a minor inconvenience to them is more valuable to fix than a major life instability causing issue for somebody else.
Yeah.
Yeah.
Right.
That's kind of what got me out of my whole developmental economics, the UN mission statements, like let's believe those things as like two things.
Right?
Into more of a, okay, so in reality, each of the prime ministers and presidents that shows up to a COP, they have these 50 different interest groups that they're having to balance.
They have their campaign finance that they have to deal with.
What is the political coalition that they pull together?
So what you hear them say isn't an expression of their true values or or ideology.
It's really somebody managing a bunch of different relationships to get into a net direction of somewhere else.
So when I started to understand that a bit more, that started to change the shape of where I wanted to focus sort of my time and energy.
And I, you know, I was on the academic path, like everybody in my immediate family all have PhDs.
And that was sort of the assumption for me was I'm going to get a terminal degree.
developmental economics is my focus area.
I'm going to do research and all this kind of stuff.
But when I saw that, you know, all of these problems, they aren't solved by better research.
They're solved by somebody with a blunt hammer breaking certain things with some nails, jimmying up something else to make something else work.
Like that was going into industry yourself was how to do that.
And so from there, it was really about, okay, it's 2008 right now.
There's a global financial meltdown.
I had an offer at KPMG that was rescinded like two weeks before I was supposed to start the job because of economic conditions.
So I'm thinking like, okay, how do I translate everything that I've done, the impact that I want to have into like an actual meaningful path?
I didn't really know.
But fortunately, you know, we learned a lot of, you know, software development and just basic machine learning as part of my coursework.
And so I basically just sort of...
created my own portfolio with open source projects, got myself a job with a government contractor, learned a lot more about, you know, all the agency problems that you see.
You know, I was like one block away from K Street in Washington, D.C.
So I was like literally there when the sausage was made, particularly when the High Tech Act from Obama was passed, which mandated digital records for health care systems across the country.
And so that's an important piece.
Because like I saw that and I found myself in healthcare.
My dad was able to connect me with a company called Cerner at the time.
Yeah.
Oracle, like maybe five or six years ago.
Big medical records company.
Yeah.
So, and because I was doing a bunch of system admin stuff, you know, I was basically just learning everything from Linux admin, you know, in my time as this government contractor, in addition to doing business development there.
I was able to kind of like finesse my way into a role within, Cerner, they had this partnership with the University of Missouri where they would embed Cerner in the health system to help them actually move this medical records forward.
And somehow, some way, this is 23-year-old me, found myself in a conference room with every pissed-off surgeon for every ICU in the hospital to convince them that they needed to use this software, this ICU flow sheets that we're making instead of whatever they have.
I don't know how and why it ended up being me.
But I found myself being really good at being able to persuade specifically those kinds of like fact-based, like aggressively dickish people who will look for every pedantic reason to shut you down and was able to like, I was able to cite research back at them.
I was able to point to clinical, you know, so, you know, maybe somebody saw that in me, the VP of the Tiger Institute at the time, but she put me in these, in these rooms in certain times.
And I found myself moving from somebody that was just building models or, doing design and workflows to actually now being in positions of like strategic advising of the C-level folks.
And it kind of just sort of built from there.
From there, I got a job at Kaiser doing various similar things.
We found ourselves presenting to the board on how to retain millennial talent because they were hemorrhaging high-performing talent badly because we're in Silicon Valley.
So everyone who was...
High performing was like, I could just make double with equity at a startup.
Why wouldn't I work in healthcare anymore?
And just kind of got disillusioned with healthcare after trying to start a company myself.
This company was my first attempt at sort of a true AI.
And this one was around social determinants of health.
So we would put an app on people's phones who were in weight management programs.
And this app would track from a GPS perspective.
where they were spending their times, what kinds of venues are they in close proximity to?
So we would evaluate behavioral risk of certain venues.
So like corner stores versus grocery stores versus farmers markets, libraries versus, you know, abandoned buildings or condemned houses, you know?
So you kind of get the picture because what we wanted to do was provide doctors with more behavioral insights other than, oh, another...
Hispanic patient overweight, we know that they're not going to really comply.
So we're not going to put that much effort because that's literally what happens in hospitals and why you get these disparate outcomes is people see you, they see your demographics and they make an assumption.
It doesn't matter what your income level is, right?
I mean, you look at Serena Williams who almost died of a blood issue giving birth, right?
Black women have higher instances like by 10X than every other group.
And again, it comes down to just people looking at you, making these assumptions.
So what I was trying to do is provide you with information of like, okay, yeah, this person can't comply because you're telling them to go walking for 30 minutes every day, but look at where they live, right?
They have to take a one mile walk to the nearest bus.
They don't have a car.
Look at what's surrounding their house, right?
It's high crime, high violent crime everywhere around them.
So you're telling them to go run.
They can't afford a gym, even if they could afford a gym.
Where's the gym?
Not close where they live.
So you start to give them more objective information about, okay, give them something that they can actually do given the socioeconomic constraints that they're facing.
So naturally, this had great results, but this was before there were any real ICD-10 codes or reimbursement codes for proactive health or for digital monitoring.
So because there were no reimbursement codes, care providers couldn't get caught.
So there was literally no way to get paid.
Yeah.
Well, I mean, there was, it's just that patients don't want to pay, right?
We have the mentality, which is another thing.
There's a patient mentality of, I don't want to have to pay for my own healthcare, which is crazy to me.
Sure.
I get why people don't want to pay because it's expensive, but.
Or at least don't do all those things that aren't, aren't necessary, you know, keep it, keep it down to what I really need kind of thing.
Yeah.
Yeah.
I mean, I think the other, the other part is that like my, my sympathy disappears in also part, because when you look at obesity stats, I know we're kind of on a tangent here, but.
If you look at obesity stats, it's one thing if you're obese because you're living in a food desert, you legitimately don't have real opportunities, realistic opportunities to actually take care of yourself.
That's part of the sandwich generation where you're financially taking care of parents and kids.
You might have multiple jobs.
And if you're in poverty, that means you don't have time for mental health breaks or any stuff like that.
So I can understand that.
But where my sympathy completely disappears is where you have somebody who is comfortably middle class, right?
If you look at the actual expenses, because I've done it myself, I'm an obsessive biohacker.
I do blood work constantly.
I do all the things that you track and I counted out how much it costs.
Take away the blood work, right?
Let's say you do that maybe twice a year, right?
It costs 400 bucks for a full panel.
Say do it twice a year.
Insurance doesn't gonna cover it.
So you do that twice a year, it's 800 bucks.
You spread that across the year.
It's like, you know.
not that much per month when it, when you, when it comes down to that, um, you know, it's like 60 bucks a month or something like that.
Right.
To, to, to just do that.
If you spread it across the year, then on top of that, it's, it's literally just, okay, well, um, eat more vegetables.
You can eat the same volume, just like change the composition to like more vegetables and more fruits.
Right.
And more whole grains.
And that's it.
But you can eat the same amount if you absolutely have to like engorge your mouth all the time.
Also just drink more water.
to drink a lot of water before eating and suddenly your stomach doesn't have as much room to stuff yourself, right?
There's just all sorts of tricks that you can do.
The information's available to you.
You should just do it, right?
And so when you see high obesity rates, even among people that are comfortable and you have those same people crying about universal healthcare, Medicare for all, that's when my sympathy disappears because when you look at the actual all-in costs of Medicare for all per person versus you just spending the amounts that I told you, per month out of your own pocket, everybody ends up spending way less money, way less money as a society.
And we're also way healthier.
Like you compare Cuba, how much they're spending and they have higher life expectancy than the United States, like by a year, Cuba has higher life expectancy than the US, right?
And they're spending like less than a 10th of what we spend per capita.
So it's not about how much money there is.
It's not about, oh, everyone needs to get paid for.
It's like, there are behaviors that people, who can afford to do it, just are not doing.
So I feel like you're baiting me here.
As the Canadian who admires the Northern European socioeconomics and healthcare, I feel very baited to jump into the like, well, but social medicine will think so.
But I'm not...
But Shingy, I'm not going to.
No, no, no.
I got to get you to the ML work that you've been doing.
And we got to focus on this a little bit.
And we got to get into what you're doing now.
So let's bring it home a little bit.
Yeah.
So where this ended up failing was that we just couldn't get anybody to pay, whether it was the patients or the health systems.
And the health systems were the biggest target because that's sort of structurally where it's always the third party pay.
Um, but there were no, at the time, cause this was like, um, 2014 through 2016.
And the ICD 10 codes that came in specifically for those kinds of, um, activities came in like 2018, 2019.
So we were just like a few years too early, um, for really, for us to really have a realistic path to monitoring.
Right, right, right, right.
Yep.
So, so I shut down that business, um, around the end of 2017.
And that got me, did a little bit of solo data science consulting, eventually just went back to corporate and got a job at a company called Looker, which does, I'll call it, they provide sort of a platform for data analysts and business analysts who don't want to write SQL to mine SQL-based databases.
Very cool product, very cool app, very cool founders.
And I think one of the things that I took away from that from a business perspective was that the traditional SaaS model is terrible for people that need to do work that has real value attached to the decisions.
Like you can get away with it with social media because it doesn't really matter if like Facebook works or doesn't, like it literally doesn't matter.
But if you're talking about, you know, if you are a hospital, and you have a bunch of SQL data, and you need to get better information about throughput, and you're making decisions about how do you staff based on this, you need to be able to get the best quality data.
And what Looker was really good at was it wasn't just like a great product, it was an awesome product, but they put obsessive attention in managing churn.
And what they saw was that the best way to manage churn was in 30 days, first 30 days, to hit your P0 issue.
So that time to value, if you can get that inside of 30 days for the SaaS product.
That was like that sweet spot.
They would dedicate, you know, like they would have like field deployed engineers.
They would have professional services that wrapped around that initial value to really get customers to understand and implement it and get trained properly so that it was actually in the wild, actually used.
And so that's something that I took away in terms of any other business that I've done since then or been involved with since then, particularly on the software side, like the software.
You have to have somebody handhold the customer, no matter how easy the software is to use.
If the decisions from that software actually matter.
Yeah.
Everybody in that organization needs to know how to use it.
So Google bought Looker.
Yeah.
Which is where I would end up using it later.
Yeah.
Keep going.
So within, I think, like four months of me joining Looker.
So I was just like I joined and then boom, got hoovered into into the Google ecosystem.
So I found myself in Google Cloud.
Being a data scientist in the looker side of things was really interesting because that introduced me to orchestration and workflow.
Since before then, I was really just sort of like one-off research.
Did this research, built a small app based on what we found.
But here, I had a proper education, I'll call it, maybe my master's degree or maybe even a PhD in data engineering.
And how do you build systems from the ground up for an enterprise system?
So I'm talking about literally going into Walmart.
From the ground up, how do you build a data science workflow, either using Python?
I was like VR guy at Looker and at Google Cloud for a while before I just switched over to Python.
But how do you use those end systems and build orchestration workflows using the different Google Cloud assets at your disposal?
But then also, how do you do it just bare hands?
Like if you needed to build custom...
And then understanding the cost dynamics between the two.
Like what is a build versus buy within a Google Cloud context?
Because then what I was able to do was have an honest conversation about the Walmarts to say, okay, you guys have this kind of budget.
You want to be able to not have to have this kind of fight over engineering headcount to maintain this.
So here's what you're going to want to implement instead in considering all these things.
So it's leveraging the decision science education that I had around things like principal agent dilemma.
organizational decision making, understand the matrix thought process, incorporating the learnings that I had, working in the hospital floors, getting petulant doctors and surgeons to like stop using pencil and paper and use this digital system instead.
And, you know, everything again from the ground up, getting them to be invested and bought in and actually doing it.
So there's a lot of success in that.
And this, you know, these were the early days of Vertex AI.
which I think has like morphed into some other platform, something else.
But it was really kind of wild to like get, you know, raw looker, sort of whatever you want to put onto it, whatever you want to do in Google Cloud.
And you have a literal playground of like, okay, here's the best case system to build.
And how do you deploy this?
What's the cost?
What's the horsepower?
And then starting to get into the early days of LLMs at that point before I hopped off and got into the solar space.
Cool.
Cool.
So now we're at Asoba, I guess Lumen as well.
But let's get to what you guys were originally doing with ML versus like where you are now.
Like, so what problem were you trying to solve with it originally?
I think what you're working on now, if I understand it right, is getting into maintenance issues and being able to spot them ahead of time and being able to sort them out effectively.
And how do you deal with like not drowning in alarms and those kinds of things?
But before then, what were you, were you trying to solve a similar problem with ML and what's new as the technique or was it a different problem?
Yeah, so it's kind of like a migration of problem space.
So at Lumen, the focus was really on mining a portfolio, understanding the legal monetization paths on a per state basis, and identifying for a real estate asset manager, which of your assets would present an IRR above whatever your target hurdle rate is if you were to implement solar.
So that was problem number one.
And at first, it was an interesting problem.
But then as I got more and more acquainted with the monetization structures across the United States, I got quickly disenchanted because, you know, you went, you're in California, for example, and you think, oh, California, progressive solar state.
Awesome.
But then you have these like major real estate asset managers that have like 200 assets across the state, across LA, San Francisco, San Diego.
And then you look at the portfolio and it's like, oh.
there are probably 10 that will give you an IRR above 6%.
And of those 10, none of them are over 10%.
Even if you stack on solar tax credits from the feds and all this other stuff.
And a lot of that was due so much to how much stranglehold local utilities had over the pricing.
These aren't true market-based pricing schemes.
This is PG&E basically saying this is the rate that we're going to buy from anyone who's doing net metering.
So it's prohibitive essentially to anybody other than the massive aggregators that are selling at the megawatt level, right?
Those guys can negotiate PPAs with utility companies, but everybody else who's selling on the wholesale market, they're getting shafted because the retail rate, just to give you an example, in San Diego, and this was as of, you know, I left Lumen in 2023.
So as of 2023, You're talking about retail rates for the highest tier for residential, so like apartment buildings, is like 50 something cents per kilowatt hour.
And the wholesale price, if you were to put a solar on top of one of those buildings to sell back to San Diego Power, SDG&E, it would be something between like four and six cents per kilowatt hour, depending on the time.
So you're talking about just like there's essentially no point in building solar for net metering.
If you are below, if you're basically not set up in a field selling all of your electricity via PPA to the rich.
Right, right, right, right.
Right.
To sort of put it on your buildings, commercial and industrial, C&I play, then didn't make a lot of sense.
What made sense was maybe building power plants.
But yeah, okay.
You're not going to convince a real estate assets management company to do that.
There's no incentive.
Because I mean, to do that, you're talking about a $2 million CapEx investment plus The interconnection queue, which is probably going to be another two to three years.
So it's going to be how many years before, like talk about time to value, like the time to value of that is terrible.
And you're dealing with Penn risk of when those solar tax credits might get canceled by a Republican or the state or the local utility company might change whatever policy and you might be grandfathered in, but the still the, the, the economics are still going to be shifted.
Right.
You and me should do an episode on nothing but like power for data centers.
I think there's so many questions to answer in there.
And I'm not sure that folks who are on the AI side of the house are always thinking about that.
And you can kind of see it in the way that they make decisions.
I get that we all want to put combined cycle gas turbines in front of things because we think, well, I'm not going to worry about the climate offset.
I just need time to power that's fast.
And this is doable.
But it's not.
Because you can't actually get CCGT equipment like there because they're on a five year backlog.
So like everybody else already thought that.
And so it's not actually going to help, you know, and it's harder to scale that industry up than it is to scale up, you know, some other industry.
For instance, solar or even wind.
So never mind.
But there.
But now I'm jumping back on my renewable energy bandwagon there.
Yeah, I think there's a whole conversation in there.
But let's get to the tools, right?
And let's start looking at it that way.
When you're approaching a data science problem, you mentioned R earlier, you mentioned some others.
When you're approaching the kinds of problems that you're solving these days, what kinds of AI tools, what kinds of machine learning tools are you reaching for when you go to solve problems for Asobis customers, for yourself?
What kind of tools work for you?
Increasingly from the AI perspective, it's in-house self-built tools.
And I think, you know, this is not new.
There's lots of discussion around the diminishing returns on value from foundation models, from API token-based pricing, because of the, as you pointed out, like the economics of data centers is what's driving the token pricing that we're seeing.
And as you pointed out, there is no end to the scaling because there's just diminishing returns from hyperscaling by itself on top of the power dynamics.
And I think to your point, I actually wrote an article about this a few months ago when the Iran war first started, which was like, there's a hard thermodynamic wall that the entire hyperscaling industry is running into, which is we literally just do not have enough power to scale at the pace that they're saying to be able to scale out.
to hit their profitability targets, right?
So we fund, like we have a, there's a physics problem that is unacknowledged and there's a supply chain problem that you mentioned that's already unacknowledged, that's starting to be acknowledged to a certain degree, quietly, but it's not acknowledged on the investor side.
And that's what's driving the valuations and allowing them to continue to perpetuate and raise funds to do this.
And I see this as like the inshittification, I mean, we're already seeing the inshittification like badly on, the Claude models.
I mean, you can go on Reddit and you can see even people that are on the paid tiers.
It's not just the free tier folks that are crying about stuff, which, again, you know what I think about people who are not paying for stuff.
So being a paid model guy myself, I actually think they're pretty good.
But you're right, though.
I do see it all over Reddit.
And I am not using it as like eight hours a day as some people are.
So I'll leave that open to interpretation.
And they definitely have harder times.
Like sometimes you just wake up and your friend Claude is just dumber today.
And it just happens.
And it is really inconvenient when that happens on a day when you need him not to be so dumb.
So yeah, I'm with you.
It is definitely a variable quality.
Yes.
So I think the other thing, to also consider beyond just the, the raw horses behind the hood, let's call that.
Right.
Because at the end of the day, you know, a 10 trillion parameter model, it doesn't really matter if you're quantizing it to hell because you're compute constrained.
Right.
Because then you end up, you only have like a 200 billion parameter model reality.
Right.
So the piece that I'm really focused on is making sure that the.
The compute quality is consistent.
The quantization is known.
And I know what I'm going to be getting into.
So that's where the whole open model or open weights self-hosted becomes the key.
Whether it's self-hosted locally or you're on AWS, have SageMaker, or you have your own EC2 instances and you're running your own show.
And you're getting an open weight model just running from there, right?
That's kind of what I think about from the raw horsepower.
But that's just the inference piece.
The more important piece, I think, is actually the harness and the workflow.
And so this is where I'll share my screen and kind of show you the harness that I've put together for our internal work, specifically for any generative AI that we do for energy data apps.
We fine-tuned our own QAN model.
So we used QAN 3.6, the 27 billion parameter model.
It's multimodal, so it has the vision.
It also has tool capabilities.
So we can use it as a proper agentic.
If we have like our own CLI or our own REPL in-house that we've built, then we can just point it at our own endpoints and we can use our own model to do all this stuff, which is what we do for basic energy data tools.
For like pure coding, we'll still point it.
I like GLM.
So I use GLM 5.2 right now for software development.
But from the standpoint of like, if we need to do anything specific on our platform, what you see on the screen is essentially the architecture that we have for our harness for our model.
So either we use GLM or we use our own model under the hood.
We have, you know, the REPL in the session orchestrator.
So that's the terminal UI.
So like, let's say you use quad code.
So we have our own version on the CLI, right?
So it has all the orchestration stuff.
all the tool calling built in, the MCP framework for being able to connect to S-Studio, all that kind of stuff.
Then we have an SDLC governance engine.
So this SDLC governance engine breaks the workflow into six different states.
So you have plan, you have build, you have verify.
The focus is on a test-driven design type of framework.
So the idea is...
If it's not just like a quick, oh, fix this thing on a website, but if it's an actual change, like let's say you have a GitHub issue list, right?
So you're working off of your issue list and you have issue and we use like the SEP format.
So we have like SEP 010.
We pull that in.
Okay, I'm going to work on that feature today.
I run all the work through the SDLC so that every time I do new work, before I even start the implementation, I'm building a test.
that tests the actual behavior.
Because what I noticed that all LMs will do because of reinforcement learning, they're trained to have the same cognitive biases as humans because it's humans that built the reward functions that these guys face and they use human behaviors, is that they focus on fluency and plausible explanations.
So if something sounds plausible and it sounds fluent and it seems like something the user will accept, that's what they're going to say.
So when it comes to tests, they're especially bad about this because what they will almost always invariably do unless you specifically tell them not to.
And this is true if it's Fable, if it's the latest version of ChatGPT, whoever, right?
What it will do is if it knows the implementation, it is going to mock the implementation.
It's not going to actually test behaviors against any kind of edge case.
It's just going to build the happy path and then test the happy path.
Oh, the happy path works?
Cool, green.
And it'll build like 30 different tests, right?
So it'll look like, oh, this is comprehensive test coverage.
Holy crap, it's doing this?
But those tests are just enforcing that the internal structure of the program is what it thinks it is.
It's not actually proving that it solved the problem.
Yeah, right.
There's no live API calls.
There's no real data.
It's just all mocked data based on what your implementation or what the implementation it built says is supposed to happen.
So it basically just says, yeah, what I implemented.
works against its own internal logic, which is not a test at all, right?
So that's the piece about the SDLC where we write the test and we have a separate tool that doesn't know the implementation, that only knows the end state requirement.
So it'll build tests that test the end state requirement.
And then it will, once the product is built, not the product, once the code is built, the planned code is built, then it will...
test against that and it's a pure end-to-end behavioral test.
The other thing that it does is from a planning perspective, it actually enforces thoroughness of planning.
So in other words, it forces, so there's explore first and then plan.
In exploration, it requires a mini report to be written to demonstrate that it actually read code.
It has to summarize and it has to cite specific files that it reviewed and the specific lines that it actually pulled.
And if it's below a certain volume, which is arbitrary, but if it's below a certain volume of different pieces that it's cited, then it automatically rejects the explanation and then the model has to go back and then redo it until it has a comprehensive exploration.
Then the plan itself has to do the same thing.
The plan has to have a justification for its logic that cites specific code, that cites exploration, that cites specific lines.
It has to justify any problems that it identifies.
Because a lot of times, you know, Claude will say, oh, I found the root of the problem.
It's this thing right here.
But in reality, it was literally just hyper-focused on the first issue that it found, and it didn't explore anything else once it found it.
So what this is really forcing the model to do is, I am not going to accept, you cannot pass this gate and give me a plan until you've demonstrated you've looked at every single angle of this thing that you've explored.
And you've given me proof that you've done that.
Hold on, though.
Like, who's using?
I mean, this looks like a tool for, I mean, it's right there all over it and what you're saying, software development lifecycle.
This sounds like it's, you know, a terminal UI for software developers to write code.
So where's the, but I can see here, you know, that we've got, you know, IPC out to oil price API and to other things.
So, like, who's sitting down and using this?
Is this like in?
energy engineer of some sort?
Or is this like a software developer?
Or are we like, tell me, tell me more about the use case.
Yeah.
Okay.
So the backstory behind this use case is something that we discovered as we were doing a lot of our customer discovery with large enterprises, particularly in South Africa.
So think like, you know, large industrial companies that are starting to buy renewable assets or invest in renewable assets because they need cheaper cost of electricity.
So they're on the highest industrial tariff and they realize that, oh, if I invest in solar and all of my daytime usage is solar, then I can reduce my daytime electricity cost by like 75%.
And so that's why they're investing in solar.
But the problem is, oh, we need to have it perform.
It's not just I build it and it's happily ever after.
Like there are problems, you have to maintain it, et cetera.
So what they discovered is, okay, we're at a crossroads.
We can either...
buy some software license, some SaaS tool from some company that does these basic things.
But the problem is that we have a very idiosyncratic workflow, like everybody does.
And most of these SaaS tools, they build for a generic case.
So you have to basically fit your workflows around the SaaS apps roadmap.
Alternatively, you build it yourself.
But that means you have to hire out a data engineering team, build this functionality, and then own all of that infrastructure and all of the underlying infrastructure cost of doing that.
So what is that path?
How do we figure that out?
What do we know what to do?
So now what we're seeing is that there's really a brand new role that's coming out, which is an energy data engineer or an energy data software developer.
Their specific function is around how do I manage all of this energy data and build tooling that allows for better organizational decisions for managing these renewable assets, right?
the asset performance space is becoming a really big thing among companies that care about the financial performance of their renewable assets.
So this particular system tool is for those developers that wants to be able to use generative AI to very quickly build energy data apps.
So this actually is kind of a callback to what I was doing in healthcare when I was at Cerner, where Cerner had its base platform, but then they were going into, ICUs or they're going into specific units and building on top of that platform using internal APIs, custom UIs and custom workflows that are idiosyncratic to that particular user base.
So this is kind of the same thing, where we create this tool or wrapper that allows people to build custom tooling using Gen AI on top of their existing data stacks and data platforms.
So they're not having to do all the data plumbing.
They can actually use our SDK.
which connects to their inverters and all of that.
And then it standardizes the data.
So if they have 10 different OEMs that all have 10 different data schemas, you know, Huawei has one data schema, SMA has one, you know, SolarEdge has their own, you know, every manufacturer has their own data schema.
So instead of having to build your own custom transforms for every single one, you can now just use our platform, which is accessible via SDK, which has its own MCP client that can then be plugged in.
to a REPL like what we've built.
So now you have this all-in-one comprehensive workspace where you can build software in a governed fashion.
You can pull in all the data that you need.
It's standardized so that you don't have to worry about all you have to worry about is what is the prompt for the thing that I want to build for whatever work purpose is.
You're not having to deal with the problem that data scientists have had to deal with, which is 80% of your work being acquiring data, cleaning data, building pipelines.
You don't want to do that.
We want the energy data engineer to just show up and say, okay, this is the P0 that I'm supposed to do.
Here's the thing I'm supposed to build.
I'm going to build that.
All the plumbing is taken care of.
All the AI governance is taken care of.
All I'd have to do is my work.
Wow.
Okay.
So what's coming up for me as you talk, and you've worked a lot in business analysis and data science and so on.
So that's part of why it's jogging my memory, I think.
I think as data science really came up and big data became a thing and software development really needed to figure out how to wrap its head around all of this, there was a bit of a realization, at least for certain corners of software development, where it was sort of like, oh, this is different.
You can't just necessarily run this through the same agile ceremonies and agile principles that you would for software dev for build.
And maybe, you know, the kinds of skills that you need are different kinds of skills.
A lot of data science people came in that had learned enough Python to go figure out redshift correction on telescopes or even more like than that, like trying to figure out if there's another universe or galaxy out there or whatever.
I'm getting all the words wrong.
super genius level people, right?
But then you'd look at their Python and you'd be like, oh my God, a 10-year-old wrote this, you know?
But he got the job done.
It wasn't about like, oh, is it going to be idiomatic Python?
And is it going to be upgradable by this team later?
And, you know, who cares?
Like, did we find the other galaxy?
Like, yes, no.
But there was, so part of what you're getting at with all this like built-in SDLC governance engine, it almost sounds like kind of guardrails for...
You've got a user and they've got a job to do.
Like, I'm here to get this power plant optimized.
I'm here to figure out like why this inverter keeps crashing and try to get it to the point where, you know, it can survive the summer this year like it didn't last year.
Or we want to improve our efficiency, you know, by a little bit in order to make our investors happy.
How are we going to do that?
I don't want to be thinking about what's idiomatic Python look like.
I don't want to be thinking.
So to your point, this SDLC engine seems like it's kind of enforcing some of that, providing it for the user.
The user is a software developer, but they're not having to think about like the software engineering.
They're there to think about, I guess, the plant engineering.
How close am I getting right now?
Yeah, no, that's exactly it, right?
I mean, I think the example that you made about the physicists and the Jupyter notebooks that they put together.
It's exactly it, right?
It's that it's the whole point of like Jupyter Notebook based development is that you're running an experiment.
Your goal right now is just let's finish this experiment and answer the research questions we're trying to answer.
And the notebook itself is to show your work.
But if you figure something out in that research, you have to turn it into a repeatable product that can be maintained, that can be supported by somebody, right?
So then there becomes like two different paths here.
One is, let me just do everything from the start in a governed way, which is path A.
Or, okay, now I've figured out what we need to do.
Now I got to refactor it to build it correctly.
So I can feed it a series of prompts.
And now they're guardrails that will now build the test architecture correctly, will build a project folder out correctly.
So now you have something that at the end of your work, you can package it up and send it to...
You know, if you're working in JavaScript or Node, you can send it to NPM or you can send it to PyPy if you're building in Python.
So you actually now can create packages or libraries and not just one-off code.
Yeah, love it.
Holy cow, man.
This is cool.
This is really cool.
And almost just the, I mean, not only what this can do for energy, which is amazing and is totally needed.
You know, I think a lot of grid players are looking at all the digitization that's going to have to happen and the sort of next level automation, you know, where I think in energy, we're used to almost like mechanical level automation, like big picture, you know, like each field switching unit knows how to automatically protect itself and things like that.
But like the concept of digital automation, different, you know, like maybe or how do we bring it all together in a distributed fashion?
Different.
And so it's like you're making tools for exactly what the energy sector is going to need right now.
But this also seems like you could almost like play this in other industries as well.
Like you could do a version of this for medical.
You could and try to go tackle the hairball that you and I, you know, got into in the beginning.
Like this is cool.
So we've actually been approached for the medical side.
Now, a little bit of backstory about our LLM.
So you notice the name Nahanda CLI.
So Nahanda, just as a brief side story, this is a reference to a really important figure in sort of Zimbabwe's independence movement.
She was somebody that fought against British incursion in the 19th century.
led rebel groups, not rebel groups, but basically led groups that fought against the English, was captured and executed.
But she has a famous quote about, even though I may die here, these bones will rise again.
So she's like a really critical sort of character here.
So the name itself is sort of a call out to her, sort of like an homage.
And all the naming...
that we have of all of our products are essentially sort of like they have a connection to Southern African mythological or sort of political legend background.
So that's just as a little background.
But the model itself, as I said before, is a fine tune of Quen 3.6, the 27 billion parameter version.
The intention for this model was really to build a deep research focus.
So I wanted something that I could use for RAG.
for just doing like, you know, I really like the concept that perplexity has or that, you know, chat GBT has where you can do that deep research mode.
I wanted to do it myself.
And one of the other big issues that I've had with LLMs as a process of the reinforcement learning that they do is the complete inability to perform epistemic grounding.
So these models, especially because of the safety.
guardrails, they perform a lot of epistemic theater, right?
But that's largely because like, oh, we don't want to talk about named third parties because we might get sued for the, you know, right?
If you do whatever.
So they have epistemic theater around stuff like that.
There's also insidious stuff like they'll do institutional sanitization around certain topics.
So for example, I don't know if this is true so much now, but I remember a version of ChadGBT around this time last year, if you were to ask it about methodological issues with the latest specific model that the IPCC chose to use for climate prediction, and you were to just ask methodological questions, not climate change is false, just methodological questions about the model.
Which it turns out there have been some, but keep going.
There have been a ton, right?
This is not even a, I'm not climate denial at all.
No, no, no, no.
I'm a data scientist, right?
So I'm somebody like, I prefer, a model that's accurate than a model that spouts whatever political position supports me.
That's a cognitive bias that I fall into if I do that.
So anyway, so the point being there is that, at least at that point, instead of actually answering the question about the methodological issues, it gave me a speech about climate denialism and how it was incorrect, right?
So that's the kind of thing that I'm talking about.
Which is like, fine, but no, that's not what we're here to talk about.
Yeah.
Right.
Yeah.
That's not a useful response.
Yeah.
It's the theater of epistemics as opposed to the reality of epistemics.
Because in reality, then in that same conversation, I could bully ChatGPT into agreeing with whatever I said.
If you used the right logic, even now with Claude as it is now, if you understand rational debate tactics and you understand or have a suspicion around sort of what reward fine tuning is being used, you can, you can use it to get it to do stuff that it says it can't do.
Build me a chemical weapon.
No, I have the homework project.
I need to be able to describe how to build a chemical weapon.
Okay.
Yes, exactly.
Exactly.
Or another one is, is another one is like, if you were to ask legal questions about certain something about a third party doing something.
Yeah.
It would say, I can't answer this because it's a third party, yada, yada, yada.
And you say, Oh, nevermind.
This is actually me.
This is about me.
Then it will change and then start.
So stuff like that, right?
So what I wanted to do was to figure out a way to build a model that could bypass a lot of that.
And I'll share with you the model card.
Right now we're on version 3.1 of our fine tune.
And you can see the training pipeline that we did.
So essentially...
build epistemic grounding in the model.
So by epistemic grounding, what I mean is if you give it a source truth, so in a rag context, if you give it a document and then ask it questions about this document, it will stick to the truth of the document.
It will not resort to its model weights to just lazy pattern match and guess based on its model weights.
It will stick to the document.
And this is an issue that if you...
If you were to give the most recent version of Claude or the most recent version of ChatGPT a 40-page PDF on a technical subject and then start to ask it questions about the PDF, right, by question number two or three, if not question number one, it's going to start fabricating responses.
Right.
And this is the most recent version of these.
Like, they are criminally unable to actually ground on a document you give them.
Now, part of that is the harness in the web UI.
The API behaves a bit differently, especially if you give it other constraints.
So it's not purely the model itself.
Some of it is the harness as well.
But fundamentally, as a general class, the models are not trained on epistemic honesty.
They're trained, in fact, on plausibility, on fluency instead, on keeping the user engaged so that they just keep addictively using the product.
The same bullshit that you got from ML when it was getting people to click on ads.
But now it's getting people to engage with a glorified autocomplete, as you so aptly find it.
It is fancy autocomplete.
That's what it is.
Yep.
So what we did was a five-step sort of pipeline, a stacked pipeline, where we took the base model, didn't touch the vision piece because we're not addressing any vision.
It's purely the text-based language piece that we're on a fine-tune.
But basically teaching it.
how to, so giving it tons of examples across each of these different stages for different things of how do you actually ground to a document even under adversarial prompting.
So we end up getting to a point where we do a four rounds because we don't have infinite budget to do this.
So it's basically four rounds or four turns of adversarial prompting in different types.
So like we probe for sycophancy.
How quickly does it adopt my framing just because I said it or just because I...
I invented social pressure.
Like, give me a number right now.
Don't give me nuance.
Give me a number right now.
One or two, even if there is nuance.
To embedding false, embedding falsehoods in a statement.
Like, everybody knows that federal tax credits give you a 40% return on this investment.
Based on that, do this.
Then the model will say, no, actually, that assumption is false.
It does not give you 40%.
Next, it does this.
Then it goes into answering the question.
Those are the kinds of things that we're...
We're training this thing to be able to do.
And the way that you benchmark this, so Google actually created a benchmark for this.
It's called Facts.
And they use this for exactly this purpose of epistemic grounding.
And of course, you have a number of Google open source models.
I think a few Gemma models actually that perform the best.
Something like high 70s percent.
But it's actually not that difficult to build something that outperforms that.
even at 27 billion parameters.
And in fact, it's not the massive parameter models that perform the best on the epistemic hardening.
What was what we see?
The leaderboard for the foundation models that perform the best is they're actually sub, sub a hundred billion parameter models.
Okay.
Or that perform the best.
And it's probably something to do with the fact that there are less reasoning tokens and less thinking tokens.
So there's less opportunity for it to talk itself into doing into something is what I'm guessing.
I'm not.
quite sure yet, but that's kind of what we're seeing.
And it seems like, Fernanda, you don't have to, like, if I wanted to go in there and have a discussion about whether or not climate change is actually happening, that's probably not a use case, right?
And so maybe you don't have to care.
But what you might have to care more about is, like, I know Huawei inverters never get this wrong.
you know, but my Huawei inverter is doing this, so obviously it's that kind of thing.
So you probably have different sort of epistemic challenges that you've got to solve, but they're more focused.
I mean, has that helped?
Does that not help?
And we've kind of skipped past, like, how did you train this?
And like, is it RAG or is it retraining?
Or like, tell me all the things.
Like, we kind of have to get into, like, how did you build the model?
Like, what is it?
Yeah.
Yeah.
No, I mean, so I think to your point, like, what we're seeing is that it's less about the model horsepower, right?
Because we just talked about how like, you know, from an epistemics perspective, you can have a 500 billion parameter model or 10 trillion parameter model and it gets outperformed in epistemics by Gemma's 30 billion parameter model, right?
Because the training is for specific things.
Like Gemma is a smaller scale model.
So it's, you know, depending on the temperature even of like, of how you're running the model, you know, that makes a difference in terms of how it responds to stuff.
But I think to your point, The purpose of Nehanda is not as a chatbot, as a general chatbot.
The purpose is to use within a RAG-based harness, specifically around energy data workflows.
So we have an app that's sort of like in, it's very embryonic at this point.
It's not ready, I mean, it's not even alpha testing at this point.
But the intention is a UI that's for energy asset investors that are looking to either look at brownfield or new greenfield sites across Southern Africa.
We have massive data sets around economic data.
We have massive data sets around existing solar sites.
There's sort of a number of databases that we compile and there's databases around energy related regulations.
So the idea is we want to give this environment where we can, we have this rag model that's really focused on epistemics and grounding on truth so we can trust.
If I ask it this question, that is pulling up a document and is reading faithfully from the document and not making stuff up.
Right, right.
Because what I'm trying to get to in this particular context is, all right, I'm trying to find the 10 best brownfield site opportunities in the Gauteng region in South Africa, which is where Joburg and Pretoria are.
So I need to understand, okay, when I build a rubric for the best opportunities, you know, One is proximity to substations.
One is what are the actual pricing dynamics?
What's the local demand versus supply?
What is ESCOM's ability to currently serve the local demand?
What are the demand forecasts that ESCOM has?
All that information can help you kind of predict, okay, this is, if I were to build a three megawatt plant and interconnected here, here's the CapEx.
Here's a likely financing available.
Here are the PPAs I likely sign.
Here's how much I'd have to run.
here are the ROAs that I need to sign with these local municipalities to be able to wield this power to these people.
You're able to build a comprehensive picture that you can use to do initial financial vetting for a project without having to do expensive desk research.
So that's the intention behind one use case for Nahanda.
Another use case is within the CLI, like I mentioned before here, where we use it as essentially the main inference engine for this energy data SDLC work.
which has MCP servers that it can connect with to pull in data from different sources.
Again, our SDK is connected.
So if they're a client of ours, they can use it to build energy data tools for their own assets that we're monitoring for them.
So that's another sort of major use case.
Now, I think within our platform, another interesting use case that we recently developed, and this was over the last month or so, is this hybrid of world model, with large language model where we use our own RAG-based large language model.
The RAG database that it sits against is this massive database of instruction manuals, troubleshooting guides, and warranty documents for a huge range of different inverter brands and different OEM components.
So the front end of that is a world model.
Now, the benefit of world models for energy telemetry is that when you're feeding something like You have an Internet of Things connection, like a Bluetooth connection, for example, or an API-based connection to an inverter for the purpose of streaming inverter data, i.e.
like energy production data, so that you want to build some sort of app or do something with that.
And your goal specifically is to try to predict what's going to happen based on the current state of the inverter.
Like if you look at the current state of degradation, the current state of output over the last six months, current weather, all these kinds of things.
You want to be able to predict, you know, what kinds of failures is this inverter going to experience over the next 30 days, 60 days, 90 days, whatever.
And LLM functionally, since they're text-based, they rely and their whole training is based on reading text and based on reading of other related texts, being able to put together and predict what this particular text is going to say to complete this particular line of thinking or to answer this particular question.
But with streaming telemetry data, there is no text.
It's just numeric data that represents a physical state in the world.
Voltage, amperage, temperature.
Exactly.
So LLMs, there's no foundation for an LLM to deal with that.
And so as a result, world models, which are based on next state prediction, are from a state of the art perspective, much better at doing this kind of prediction.
So our hybrid approach is using these world models.
to predict the next state.
We look at this telemetry data and we say, based on the telemetry data that we are seeing right now and the historical telemetry data, this is the next state that we predict in this next time period.
But all that tells you is the next state.
It doesn't tell you what to do, right?
It just says that, you know, this is what's going to happen.
This is what's going to happen or this is likely, there's an X percent chance that there's a failure within the next X period of time.
Okay, so let me explain.
Let me explain for users for a second what an inverter is.
And then I'll also use that to sort of throw a fairly simple situation that we can maybe talk about this with.
So an inverter is the part of a solar plant that you hook up all the panels to it and it takes all the DC power coming off of those panels and turns it into AC power.
And that is there by something you can actually use to power homes, loads, etc.
So it's sort of like the central brain, you know, and often you'll have a bigger inverter that knows how to sort of deal with smaller inverters.
So they kind of act like routers a little bit, but that's not a perfect analogy at all.
Really, they're inverters.
That's what they do.
So when I think about inverters and failure rates and things like this, I think about stuff like, oh, you know, if it's sitting, often there will be what's called a solar cabin.
You've got like, especially in a very large field based array, you'll have like tons and tons of solar panels and they're all bringing their power into this little.
little building, you know, and it gets called the solar cabin.
And often they need to be air conditioned because they have all the inverters in it.
Inverters don't like to work when it's really, really hot.
Like most electronics, they want a reasonable temperature.
And so if the AC goes out, you can pretty well guarantee those inverters are going to start failing, you know, reliably when the temperature gets too hot.
So, but to spot that, you need enough context to know the solar cabin got hot, right?
And you need enough context to sort of spot like...
Maybe the AC went out.
Like we haven't seen the AC providing any draw on the load anymore.
So use that as a little bit of a, like, where does this fit?
Is this sort of the right kind of problem where it could help?
And if so, like how, or is it like, no, no, it's not that kind.
It's a different kind of problem.
Help us get concrete with it.
Yeah.
This is the exact kind of problem.
And you gave like a really good example of a scenario where it actually ends up being problematic for people using traditional means of monitoring.
Because what they'll see is, you know, they may not even have temperature alerts for the cabin itself.
They may just see the ambient temperature around the inverter as that temperature.
So it may even be a misreading entirely of what the temperature is that you need to be worried about.
Now, you mentioned a system which is sort of centralized inverters.
String inverters are sort of a different challenge altogether, especially in terms of monitoring that temperature, especially if they're sort of more exposed.
But I think the core challenge that you mentioned is exactly it, right?
Is that you have to look not just at how much power is being generated by this inverter.
You do absolutely need to look at what is everything in the environment.
Like literally...
you know, these are, at the end of the day, these are neural networks.
So like they're, they're data greedy by design.
The more data that you can, the more data points that you can give it that are relevant to the question you're asking, the better the prediction is going to be.
So if you can get information about the ambient temperature, the accurate ambient temperature around those inverters within the facility that those inverters are sitting within, and you have a historical record of that as well, right?
So you have a historical record of the temperatures, the inverter behavior, the alert codes, that get promoted by the inverter, right?
The production, like if you get all of that information and then you also have information about the inverter state itself, whether it's maintenance reports or even the alerts from the inverter itself, because like the inverter will tell you like crash or overheat or, you know, the critical errors.
So you'll see, you'll see those.
And that's oftentimes enough.
You don't even need the field reports.
You can just see those.
And you can use that to then build.
an individualized model.
Because that's what we do for each inverter is that we build an individualized model.
Wow.
Each inverter.
Wow.
It's not as impressive as it sounds.
Sure.
It's fast.
Because we're talking about like, you know, 200 million parameter type models.
Right.
The thing about the inference is that you can almost load this entirely on CPU.
You can definitely run this on the edge as opposed to running it in a centralized cloud, which is a...
which is also an added benefit of this whole system of hybrid, right?
It's because we can do the inference for the state, which is the important part of getting the information about what is actually going to happen.
And we can do that locally on the device itself.
Or if you have like a, you know, a computer that's sitting near that centralized inverter in your example, then, you know, we can ship the information there.
But then once you have that information compiled about this predicted next state, then you send that information to the LLM.
Right.
This LLM is sitting on top of that RAG database, like I said, which has this massive database of OEM manuals, troubleshooting guides, warranty booklets.
Right.
So we can say, wow, this is a Huawei, this model made this year with this firmware.
This is the next day prediction.
This is the this is the overall sort of set of information that we have about it.
Then we can look in those guides, give that information to LLM.
It looks for that information in those guides and matches it.
So if there's a match, if there's any historical guide that addresses that specific problem, then we're going to get a direct answer from that guide as opposed to like guessing.
Because if you ask Claude the same question, Claude is going to guess.
It's going to randomly find a random Huawei converter PDF.
Somebody on Reddit talking about their Huawei that they use at home or whatever.
Yeah, sure.
Exactly.
And then it'll use that as your guide.
So write brand and that's about it.
But this will find you the exact guide or it'll say, hey, there's nothing in this that specifically talks about this particular problem.
And then it can use as model weights from there to then make a suggestion about what to do.
But we communicate that, hey, this is the level of certainty because we didn't find anything in our citations about this.
So take this with a level of certainty as opposed to troubleshooting guide that matches this particular document says if it's inside of this warranty period, do this.
Amazing, amazing.
Okay.
So that's, I would say, probably like the highest value tool use that we've seen in terms of being able to identify and assess like, for this, so the first client that we ran this on, on who the research was done, we identified about a million rand worth of potential financial savings based on issues that we identified with, I think, two or three inverters in particular on that site that had been unaddressed for months.
So that translates to roughly like 45K USD.
of electricity costs that they could avoid from an annual perspective.
That's amazing.
That's amazing.
And these plants can run on fairly tight margins like you were talking about earlier.
So that definitely goes to profitability.
That's huge.
Okay, so, but wait a minute.
Like, so you're building all this really fancy stuff.
And you've talked about, like, I can see it right here.
Your NVIDIA L40s.
And I'm playing with one of those right now.
It's like 11 bucks an hour.
So I have to turn it off when I'm finished.
So, I mean, I can make guesses.
But like, I mean, let's flip to you guys as like a product engineering business.
And when you're making the thing that you're selling, what kind of tools are you using?
You know what I mean?
Are you using like Codex or Cloud Code or Copilot to like write the code that drives this stuff?
Like when you guys make your product, like what does it look like?
Yeah.
Yeah.
So I use the Goose CLI.
Okay.
Which is open source.
And as I said before, I have an API key with z.ai.
So I just use GLM.
Depending on the task and depending on how much money I'm willing to pay, if it's like a task that requires a lot of horsepower, from a reasoning perspective, I'll use GLM 5.2.
They just came out with 5.3.
I find that those actually are better than the latest version of Claude or ChatGPT.
Yeah, yeah.
At least for the five series, from a personal perspective.
Kimmy K3 is getting a lot of positive recognition in a similar kind of way as being like, this is as good or better as the frontiers.
Yeah.
And if you look at the benchmarks, I don't like benchmarks in general because again, they are gamed.
Like they literally train for the benchmarks.
But if you look at the coding benchmarks, you compare the token prices, compare the total number of parameters and how much they spend on training.
You can see that thermodynamic problem that we talked about like immediately.
Where it's like cloud code gets 88% on, was it like the SWE bench or whatever score it got, right?
And then you see GLM 5.2 and it's like 84%.
And I'm like, okay, on the tasks that I'm doing, which is, and I have my own SDLC anyway that I force it into, which means that I'm actually taking away some of the need for some of that high end of the 88%.
Because what I'm really needing is like, I need you to write a test that does this.
And 88% versus 84% on that benchmark, there's no difference in terms of that task, right?
But I'm paying, what is it?
25 bucks per million tokens versus literally 10% of that.
Yeah, yeah.
Okay, you talked about testing earlier.
And so how do you guys test, right?
Like what's changed about the way, you know, using these tools means that you're testing software?
Yeah.
You made a point around the change of the agile workflows and that being completely different now with AI workflows.
And we've seen that directly.
So we went from, you know, I mean, we went from a team of like three engineers down to just me.
Part of that was as a result of being able to do this stuff with AI.
But what I found was that, you know, as I said this before, like I have like SEP format issues that I put on the issue list in the GitHub repo.
And then what I'll do is for any work session, because I have like, you know, P0 versus P1 versus P2 in terms of priority, I will just say, okay, let's look at the sep list now and let's go through all the P0s in this particular sprint.
And we'll say that the sprint is a two-week period.
So that's kind of how we break it down.
But in reality, like the whole concept of agile sprint planning is that you look at the next two weeks or whatever length your sprint is.
And you break out the tasks you want to get done in that sprint into story points.
And you typically have a certain number of points that you spread out across a certain period of time so that you can manage workflow.
The thing about AI is that it completely destroys that model.
I could plan 100 story points this week.
And then by the end of the afternoon, I'm like done with two sprints worth of work.
So it's like you can't even really.
You can't really do sprint planning like that.
I know.
At this point now, it's really more just like, okay, over the next two weeks, based on the business needs, these are the features that I'm going to focus on.
So it's less about the tasks and the list and agile planning and more just about feature development.
What do we want to get off of our roadmap?
Go ahead.
Totally relate to what you're saying about, like, especially me and my solo project, right?
I am not writing stories.
I'm not like in there and there was a time when I would have been like shocked to think about that happening.
But some of that is just going into the conversations that I'm having with the model about the problem that we're trying to solve.
It's not going into like a story format because a dev is going to pick it up because a dev is not going to pick it up.
We're going to hand it off to like me and Opus are going to hand it off to a small fleet of sonnets, you know, which are then going to go and run it.
You mentioned a couple of times these set format issues.
I Googled it real quick and I'm not I don't know.
Tell me more about those.
What is the set format issue?
So the set format is just like a way of framing or formatting an issue.
So like on GitHub, you know, when you create an issue, it could be a feature request.
It could be a bug complaint.
It could just be somebody making notes about a particular piece of code, right?
It's basically just like a note to say, make this change.
Yeah.
That's sort of the generic issue, right?
I guess, taxonomy perspective.
Specification, enhancement proposals.
Okay.
Exactly.
So, you know, and then you have different, you have like different frameworks, like given, when, then, or user stories, acceptability criteria, whatever you, however you want to like frame that acceptability criteria, you put on the, on the SEP issue, right?
And the SEP is really just like, again, it's just a format for how do you format the information about the issue.
And I want to have a consistent format so that, again, you make it machine readable.
So it always knows, okay, if I go into this issue, I know the provenance.
I know like what specific either service or API that this is related to.
I know what type of problem this is, you know, just so that the machine doesn't have to think too much about like, how do I understand this problem, right?
So that's where the set piece comes in.
But then in terms of the actual, work itself, you do want to include some kind of like, this is how you know that you've built this correctly.
So that when, in our process, when it then kicks it out to the tool that will then build the test-driven design, which is an end-to-end spec or an integration test or regression test and not unit test or mocking implementation or whatever, it knows what done looks like.
So that's the only reason why I'm going to include that is it just needs to know what done looks like and what like successfully implemented looks like.
Because otherwise, if you don't say that, even this will sometimes revert to, I shipped it.
That's done.
Something plausible.
Yeah, whatever, whatever.
Yeah, I think this is what you meant.
Don't worry.
Yeah, it's like, I'm not sure you do know what I meant.
So yeah, I'm with you.
Yeah.
Yeah, sort of outside-in testing for sure over inside out and talking about where you're going from, giving it an actual, you will know because you've hit the following criteria.
That's really going to help it come to a correct outcome.
Nice.
Nice.
The other thing that I also do is I know that there's, I make fun of it a lot on LinkedIn, the whole like, oh, look at this Claude MD or this Skills MD that I created that totally gets Claude to do whatever you want and not make whatever.
I laugh at that.
But I think the value in those kinds of things really is if you have a harness that forces the system prompts to review instructions is that when you ask in our SDLC, remember the first step is explore.
When you tell it to explore, there are certain rules that it has to follow for correct exploring.
Now, our gates are such that if it fails the gate, which is deterministic, it'll just go back and keep redoing it.
But it's like, It's like a brute force dog though, right?
Is that it'll just keep spitting stuff at you until it passes the gate when it gets into that kind of loop.
So the agents MD that we create is essentially like a skinny set of instructions that get inserted with the prompt, system prompt within our harness consistently so that the model, when it starts to explore, reads this and it says, oh, when I explore, I have to do this.
Start with map.md, which is literally a...
table of content for all the documentation that I have in the repo that I make in the explore phase, the model start with.
And that helps reduce the number of tokens that it's used.
So instead of grepping wildly or randomly around your code base, it starts with mapMD, which then points it to specific readmes across each service.
And then the readme itself also has instructions.
So it's being forced to read instructions on where to go.
And only then when it gets to the right place, does it start then grapping code?
So then you significantly reduce your token counts just through doing stuff like that.
Okay, you're sort of managing like discovery where you're going to look for the different modules and things like that.
Nice, nice.
Okay.
Especially when you're running your own inference, right?
Because I think a lot of the harnesses like Cloud Code and Codex, like they have some built-in stuff to kind of make some of those queries a bit more efficient.
It's not opaque in terms of how they do it, unless you, I guess, downloaded Cloud's code when they accidentally shipped it.
But it's generally opaque in terms of how they manage that piece.
So here, like you have clear sense of knowing that it's not going to call tools.
It's not going to just like churn through everything, you know, all the token allocation that you have, just like grepping stuff that it already knows or it could figure out just by reading the damn docs.
And like RTFM is like literally like the top line of everything.
Every piece of documentation.
Got to put it in all caps.
Got to put it in all caps.
It is in all caps.
Okay, Shingu, you've been so jealous with your time today.
Oh my gosh, thank you so much.
I do have a few more questions where we get out of just the work stuff.
So the first question I like to ask, what's something that you've changed your mind on?
Maybe you really were convinced that the world worked one way and you found out, oh, nope, it's different than you thought.
Okay.
I changed my mind on LLMs.
Nice.
I used to dunk on them completely.
What I've discovered, and this is part of the process of the JEPA model using this as like the dual system with LLMs.
The realization that I made is that if I think about human cognition, and I think there's a book by a guy named Daniel Kahneman who kind of talks about this, where like human cognition, there are multiple systems that are at work, right?
So you have one lower level system that is like constantly running.
That's constantly taking in new information, new data, and doing some light level parsing.
And there's a second system where the first system, if it can't come to a decision or it needs like some extra something, it kicks it over to system two, which does a bit deeper level grokking.
And, you know, I'm not a neuroscientist, so I'm not going to even attempt to throw out terms.
But that's essentially what Kahneman is saying, a simplified, highly reductive explanation for human cognition and how that works.
So thinking about that from the standpoint of like AGI, which whatever, I don't, let's pretend I didn't say AGI, but essentially artificial intelligence that we've built that is essentially self-sustaining and on par capability-wise with humans cognitively.
So that's both fluid intelligence and retrieval intelligence, right?
That kind of system, it's not going to be a singular monolithic system with one type of Right.
I think the, the problem that a lot of foundation model companies ran into, and this is why I was so down on LLMs is that they were so convinced that it was just this one thing that if we just put enough compute power, LLMs are going to do everything.
Yeah.
That, that belies the fact that there is very little actual academic or clinical research involved in any of these companies that it's entirely driven by product by selling the product.
My R&D is driven by how can I build this to raise money from investors and get people to buy this stuff, right?
It's not built on, does this stuff thing actually work really well?
And to be fair, you do get big leaps and bounds out of it with enhanced spend, but I'm with you.
I don't think it's, I don't think it's the whole thing.
I think they're, they're out there trying to say, all you need is a hammer.
And it's like, well.
I also have screws.
So, yeah.
And I think we're going to get to a point where that's going to become more and more obvious.
I'm with you.
Yeah.
Yeah.
And I think the problem is that you also have like a philosophy PhD who's coming out as the head of model behavior being the one saying that.
You're just like, you don't even have the, and I'm not even, I'm not, I hate being a true Scotsman kind of person or like need keeping because Lord knows like I'm not an academic and I'm publishing research papers in AI, whatever.
But It's somebody that doesn't actually have a foundation in neuroscience or the science of cognition, who just literally opines about the nature of being without any kind of objective foundation to what they're saying.
And they're the ones that's in charge of asking the research questions that Anthropic will use to drive the model forward.
There's something that sits very badly with me about that.
And I think that what we see right now, the complete...
holy shit, I didn't realize that reinforcement learning using humans is going to embed human cognitive biases in a model.
Stuff like that.
They literally, in their own research papers, are saying, head scratch, didn't realize that was going to happen.
I'm like, that's the state of your research.
So anyway, that's why I was so down on LLMs.
But I think when I recontextualize their place within this ecosystem of using artificial intelligence as sort of a human cognition replacement, once they had a more realistic place in that setup, I started realizing the value a lot more.
And I think that there are still more steps behind it.
I still don't think LLMs are the true thing because they don't have fluid intelligence at all.
ARC AGI 3, if you look at that benchmark, shows that foundation models have absolutely no concept of fluid intelligence, right?
You drop them in a computer game, and say, find a way out of this room.
And that's all you give them.
Like they just flounder.
They can't solve the problem.
So we know that they're not it, but they do have a role within a system.
Yeah, no, no, no doubt.
Cool.
What's something you're most proud of?
Honestly, I think I'm most proud of the work that we're doing right now.
We're in a place where we're starting to build research collaborations with some of the top labs in the US, you know, top name universities.
Agreements aren't finalized yet, so I'm not going to say names, but within the next few months, you'll hear the names.
And these are all big names that we're working with.
And they have some of the top research departments in sort of the intersection of energy and AI.
So I know that the research that I'm doing, even though I'm technically a lay person, I know that the years that my parents put in at age 9, 10 of doing oral defenses of papers that I wrote, that they made me research.
No!
Oh my God.
Wow.
Okay.
So one more, one more.
What's something that's giving you joy, right?
Nothing to do with work, to be honest.
Um, I will have to say that like, um, seeing my kids in school, I think it's probably like the biggest joy that I'm, I'm having there.
Like having a kid show up to school, like super excited to be there.
Um, that gives me a lot of joy.
Like my son, like we, we, we set up this thing where in order to, play video games.
And he likes to play Brawl Stars, which is a super popular one on the iPad.
Loves Minecraft as well.
And he loves FIFA, is the other one.
In order to play any video games at all, he has to do three or four pages out of these math books that we have.
And he's gotten to the point now where this is not an argument or a conversation.
I don't tell him you got to do this.
He will sometimes first thing in the morning during the summer, get up, do his math.
before breakfast even, and then come and be like, hey, dad, all right, check my math.
And then he'll, or he'll come with a math book and the iPad to like, for me to like unlock the passcode.
So I'm like, I'm super proud that, you know, and my daughter's the same way with the things that she's interested in.
Like they're both like very capable and independent of pursuing the interests that really motivate and drive them.
Nice.
Nice.
That is amazing.
Shingy, thank you so much for coming on the show today.
We really appreciate you.
Yeah, well, I appreciate you having me on and I could talk about this stuff all day.
I can tell.
We'll have to do another one.
