

Osiflow HMI: the product
The model proposes. Code decides what's real.
AI Product Design
Physical AI Software
Context
An AI-native audit engine where every recommendation carries its evidence and nothing ships without a reason attached.
Context
HMI teams across automotive, robotics, industrial, and medical-device work kept describing the same trust gap. Nobody will sign off on an AI recommendation they cannot trace back to a reason.

Highlights
That gap does not disappear when a tool becomes faster. It closes only when the recommendation can show what it knows, how it knows it, and where human judgement still begins.
The Challenge
Every AI design tool I looked at, including early versions of my own, produced output with total confidence and no receipts. It would tell a designer something was wrong or something should change and offer nothing underneath that claim.
No standard cited, no assumption named, and no way to check the work. Confidence with nothing behind it just means nobody checked.
Solution
AI Nominates, Code Certifies
Here is the idea in plain terms, because it deserves to be understood, not just believed.
The AI's job is to nominate. It looks at a design and proposes what might be wrong or what could improve. It is allowed to have opinions, and it is allowed to be wrong sometimes, the way any capable colleague is.
What it is never allowed to do is grade its own homework. Every recommendation passes through separate code that checks it against the standard it claims to cite. If the AI cannot point to a real rule and a real reason, the recommendation does not get to sound certain. It has to say, plainly, that it is not sure.
Think of it as a junior colleague who drafts a proposal and a senior colleague who has to sign it before it goes anywhere. The AI is the junior colleague: fast, generative, occasionally overreaching. The code checking its work is the senior one: slower, unglamorous, and the reason anyone can trust what finally reaches the screen. Neither works alone, and that is the whole point.
Outcome
The product now shows how a recommendation earned its confidence before asking a person to trust it.


Where It Started
Confidence without evidence was not trust.
AI design tools, including early versions of my own, produced confident recommendations with no receipts. Designers could be told something was wrong without seeing a standard, an assumption, or a way to check the claim.
The product needed an answer that could prove why it deserved belief.

The Verification System
The model proposes, but evidence earns the verdict.
The AI nominates what might be wrong, while separate code checks measurable claims against cited standards. Anything unproven must expose its assumptions and lower its confidence.
Carbon provides the component foundation, while Claude and its subagents accelerate work that can be verified.

Judgment Log
The system had to know when not to sound certain.
The model could not grade its own homework. Deterministic detectors verify measurable claims, experienced judgement handles what cannot become a rule, and human review remains the final gate. I kept the expensive product decisions and delegated repeatable execution that could be checked.

Outcome & Closing
Trust became a visible part of the interface.
The beta demonstrates a working evidence model rather than a promise about responsible AI. Each recommendation exposes its source, assumption, confidence, and verification state before acceptance. It closes on the same principle used to build it: speed matters only when the work remains precise and accountable.
More Works
©2026
FAQ
01
What kind of problem is worth bringing to you?
02
Do you work as a designer, founder, or advisor?
03
What do you need before a first conversation?


Osiflow HMI: the product
The model proposes. Code decides what's real.
AI Product Design
Physical AI Software
Context
An AI-native audit engine where every recommendation carries its evidence and nothing ships without a reason attached.
Context
HMI teams across automotive, robotics, industrial, and medical-device work kept describing the same trust gap. Nobody will sign off on an AI recommendation they cannot trace back to a reason.

Highlights
That gap does not disappear when a tool becomes faster. It closes only when the recommendation can show what it knows, how it knows it, and where human judgement still begins.
The Challenge
Every AI design tool I looked at, including early versions of my own, produced output with total confidence and no receipts. It would tell a designer something was wrong or something should change and offer nothing underneath that claim.
No standard cited, no assumption named, and no way to check the work. Confidence with nothing behind it just means nobody checked.
Solution
AI Nominates, Code Certifies
Here is the idea in plain terms, because it deserves to be understood, not just believed.
The AI's job is to nominate. It looks at a design and proposes what might be wrong or what could improve. It is allowed to have opinions, and it is allowed to be wrong sometimes, the way any capable colleague is.
What it is never allowed to do is grade its own homework. Every recommendation passes through separate code that checks it against the standard it claims to cite. If the AI cannot point to a real rule and a real reason, the recommendation does not get to sound certain. It has to say, plainly, that it is not sure.
Think of it as a junior colleague who drafts a proposal and a senior colleague who has to sign it before it goes anywhere. The AI is the junior colleague: fast, generative, occasionally overreaching. The code checking its work is the senior one: slower, unglamorous, and the reason anyone can trust what finally reaches the screen. Neither works alone, and that is the whole point.
Outcome
The product now shows how a recommendation earned its confidence before asking a person to trust it.


Where It Started
Confidence without evidence was not trust.
AI design tools, including early versions of my own, produced confident recommendations with no receipts. Designers could be told something was wrong without seeing a standard, an assumption, or a way to check the claim.
The product needed an answer that could prove why it deserved belief.

The Verification System
The model proposes, but evidence earns the verdict.
The AI nominates what might be wrong, while separate code checks measurable claims against cited standards. Anything unproven must expose its assumptions and lower its confidence.
Carbon provides the component foundation, while Claude and its subagents accelerate work that can be verified.

Judgment Log
The system had to know when not to sound certain.
The model could not grade its own homework. Deterministic detectors verify measurable claims, experienced judgement handles what cannot become a rule, and human review remains the final gate. I kept the expensive product decisions and delegated repeatable execution that could be checked.

Outcome & Closing
Trust became a visible part of the interface.
The beta demonstrates a working evidence model rather than a promise about responsible AI. Each recommendation exposes its source, assumption, confidence, and verification state before acceptance. It closes on the same principle used to build it: speed matters only when the work remains precise and accountable.
More Works
©2026
FAQ
01
What kind of problem is worth bringing to you?
02
Do you work as a designer, founder, or advisor?
03
What do you need before a first conversation?


Osiflow HMI: the product
The model proposes. Code decides what's real.
AI Product Design
Physical AI Software
Context
An AI-native audit engine where every recommendation carries its evidence and nothing ships without a reason attached.
Context
HMI teams across automotive, robotics, industrial, and medical-device work kept describing the same trust gap. Nobody will sign off on an AI recommendation they cannot trace back to a reason.

Highlights
That gap does not disappear when a tool becomes faster. It closes only when the recommendation can show what it knows, how it knows it, and where human judgement still begins.
The Challenge
Every AI design tool I looked at, including early versions of my own, produced output with total confidence and no receipts. It would tell a designer something was wrong or something should change and offer nothing underneath that claim.
No standard cited, no assumption named, and no way to check the work. Confidence with nothing behind it just means nobody checked.
Solution
AI Nominates, Code Certifies
Here is the idea in plain terms, because it deserves to be understood, not just believed.
The AI's job is to nominate. It looks at a design and proposes what might be wrong or what could improve. It is allowed to have opinions, and it is allowed to be wrong sometimes, the way any capable colleague is.
What it is never allowed to do is grade its own homework. Every recommendation passes through separate code that checks it against the standard it claims to cite. If the AI cannot point to a real rule and a real reason, the recommendation does not get to sound certain. It has to say, plainly, that it is not sure.
Think of it as a junior colleague who drafts a proposal and a senior colleague who has to sign it before it goes anywhere. The AI is the junior colleague: fast, generative, occasionally overreaching. The code checking its work is the senior one: slower, unglamorous, and the reason anyone can trust what finally reaches the screen. Neither works alone, and that is the whole point.
Outcome
The product now shows how a recommendation earned its confidence before asking a person to trust it.


Where It Started
Confidence without evidence was not trust.
AI design tools, including early versions of my own, produced confident recommendations with no receipts. Designers could be told something was wrong without seeing a standard, an assumption, or a way to check the claim.
The product needed an answer that could prove why it deserved belief.

The Verification System
The model proposes, but evidence earns the verdict.
The AI nominates what might be wrong, while separate code checks measurable claims against cited standards. Anything unproven must expose its assumptions and lower its confidence.
Carbon provides the component foundation, while Claude and its subagents accelerate work that can be verified.

Judgment Log
The system had to know when not to sound certain.
The model could not grade its own homework. Deterministic detectors verify measurable claims, experienced judgement handles what cannot become a rule, and human review remains the final gate. I kept the expensive product decisions and delegated repeatable execution that could be checked.

Outcome & Closing
Trust became a visible part of the interface.
The beta demonstrates a working evidence model rather than a promise about responsible AI. Each recommendation exposes its source, assumption, confidence, and verification state before acceptance. It closes on the same principle used to build it: speed matters only when the work remains precise and accountable.
More Works
©2026
FAQ
What kind of problem is worth bringing to you?
Do you work as a designer, founder, or advisor?
What do you need before a first conversation?

