Measure, do not imply
Versions, test results, failures, and promotion decisions stay visible.
Cardona CT Lab / Living Research
Self-hosted intelligence, measured in public.
PlantForge, CodeForge, ImageForge, and MeshForge are applied local intelligence tracks. This whitepaper records what each system can do, what it cannot do, and the evidence required before capability expands.
Research snapshot
Useful local intelligence can grow without quietly expanding authority. Every improvement must be visible in repeatable evidence.
Versions, test results, failures, and promotion decisions stay visible.
Models explain or propose. Deterministic systems and people approve consequential action.
A newer or larger model is a challenger until it passes the defined gate without regression.
Current state
Loading the model registry.
ImageForge visual lab
Reference direction, local baselines, and trained candidates remain visibly distinct. A good reference is not evidence that the local model produced it.
Publication boundary: only approved, public-safe outputs appear here. Private prompts, evaluator rubrics, rejected generations, and training assets remain outside the site.
MeshForge mesh lab
The first Stable Fast 3D runtime completed locally. Its artifact reopened correctly, but topology quality blocked promotion.
Candidate 01
Stable Fast 3D produced a reopenable GLB in 70.89 seconds at 6171.84 MiB peak VRAM. The artifact contained 4,136 non-manifold edges, so it remains evidence rather than a promoted result.
Silhouette, completeness, scale, and manifold checks.
Non-manifold edges, intersections, density, and repair distance.
UV integrity, texture coverage, and visible seams.
Human-approved GLB or OBJ only. CAD claims require a separate parametric track.
Operating boundary: MeshForge may generate review candidates. It cannot approve topology, overwrite source assets, or publish an export.
Capability register
"Current" means selected for its bounded role. It does not mean autonomous, generally capable, or safe outside that role.
| Project | Model | Base | Role | Evaluation | System evidence | Decision |
|---|---|---|---|---|---|---|
| Loading model evidence. | ||||||
Measured growth
Bars show comparable suite pass rates. Missing measurements remain unmeasured, not zero.
Loading benchmark history.
Public capability lab
Curated demonstrations show how implementation quality changes across releases. They are not private evaluator cases or hidden-test outputs.
Publication boundary: prompts, judges, reference implementations, and raw candidates from sealed evaluations remain private. These specimens are purpose-built for explanation.
Decision record
Every retained, rejected, or testing decision carries a reason and an evidence class.
Loading promotion decisions.
Promotion method
Capability is promoted one bounded role at a time. A model never inherits authority from a good demo.
Name the task, inputs, outputs, and forbidden actions.
Use deterministic tests, hostile probes, and untouched holdouts.
Retain evidence without changing source, hardware, or production state.
Measure useful completion, missed requirements, regressions, and unnecessary change.
Advance only when the full gate passes and the authority boundary stays intact.
Operating boundary
PlantForge cannot water a plant. CodeForge cannot commit, publish, or deploy. ImageForge cannot select or publish its own output. MeshForge cannot approve or export a generated asset. These are structural constraints, not promises in a prompt.
PlantOS validates evidence and owns care-state decisions.
No actuation routeDisposable worktrees and authenticated evidence contain model output.
Human commit requiredVersioned workflows and frozen briefs produce candidates for review.
Human selection requiredLocal image-to-mesh systems produce bounded geometry candidates.
Human export requiredEvaluation history
Loading the evidence timeline.