Measure, do not imply
Versions, test results, failures, and promotion decisions stay visible.
Cardona CT Lab / Living Research
Self-hosted intelligence, measured in public.
PlantForge, CodeForge, ImageForge, MeshForge, and VisionForge are applied local intelligence tracks. A currently unbranded audio research track is planned. This whitepaper records what each system can do, what it cannot do, and the evidence required before capability expands.
Research snapshot
Useful local intelligence can grow without quietly expanding authority. Every improvement must be visible in repeatable evidence.
Versions, test results, failures, and promotion decisions stay visible.
Models explain or propose. Deterministic systems and people approve consequential action.
A newer or larger model is a challenger until it passes the defined gate without regression.
Current state
Loading the model registry.
VisionForge visual QA lab
On six private cases, Qwen 7B scored 1/6 and Gemma 12B scored 2/6 model-only. Both missed overlap and aspect-ratio defects and produced false positives on the clean control. Qwen3.5 27B scored 0/2 after both cases timed out. No model is promoted.
A local image enters a loopback-only analysis route.
The model reports visible evidence and labels inference separately.
Deterministic code owns geometry, luminance screens, duplicate checks, schema, and authority.
A person decides whether the finding is correct and actionable.
Operating boundary: VisionForge may identify possible aspect-ratio, clipping, composition, overlap, text-legibility, artifact, and PII risks. It cannot declare compliance, identify a person as fact, change the source image, or approve publication.
Human review contract: Every finding requires an accept, reject, or revise decision. The completed review binds to the exact report digest, preserves the original evidence, and grants no publication or source-modification authority.
Local review workbench: The browser UI reads selected reports and images locally, overlays supplied evidence regions, exposes provenance and confidence, and keeps export disabled until every finding has a decision and rationale. Automated browser checks verified structure, empty state, disabled export, console health, and narrow-screen fit. Local-file injection was not exercised by browser automation.
Kimi K3 may propose difficult cases and critique approved outputs using public, synthetic, or explicitly redacted material. Its contract rejects private input labels, sealed-answer exposure, external promotion authority, and results that do not require independent verification.
Authority: advisory only. It cannot promote a model, change an evaluator, modify source files, publish, deploy, or convert its own claim into verified evidence.
ImageForge visual lab
Reference direction, local baselines, and bounded refinement studies remain visibly distinct. A good reference is not evidence that the local model produced it.
Publication boundary: only approved, public-safe outputs appear here. Private prompts, evaluator rubrics, rejected generations, and training assets remain outside the site.
Audio research / planned
Planned. No base model has been selected and no audio capability is claimed. The first bounded use may be one optional Falling Blocks interface sound after every safety gate passes.
Define an original audio task, excluded references, and accessibility requirements.
A local model proposes a loop or sound effect with seed and settings recorded.
Tools check decoding, duration, silence, clipping, loudness, and loop seams.
A person reviews musical fit, originality risk, licensing context, and publication.
Operating boundary: no voice cloning, identity impersonation, living-artist imitation, autonomous licensing decision, autoplay, or autonomous publication. No model is selected and no audio has been generated.
MeshForge mesh lab
The first Stable Fast 3D runtime completed locally. A verified ImageForge artifact may guide visual direction, but it is not geometry evidence and topology quality still blocks promotion.
Candidate 01
Stable Fast 3D produced a reopenable GLB in 70.89 seconds at 6171.84 MiB peak VRAM. Static inspection found 12,146 vertices, 19,984 triangles, 169 connected components, and 4,136 boundary edges. UVs and two packed textures passed, but topology kept the result at Runnable.
Silhouette, completeness, scale, and manifold checks.
Non-manifold edges, intersections, density, and repair distance.
UV integrity, texture coverage, and visible seams.
Human-approved GLB or OBJ only. CAD claims require a separate parametric track.
Input contract: geometry evaluation starts only after a person approves a controlled reference set with front, side, and three-quarter views of the same object, consistent lighting, a plain background, visible scale, and no cropped silhouette. A composed editorial image cannot satisfy this gate.
Operating boundary: MeshForge may generate review candidates. It cannot approve topology, overwrite source assets, or publish an export.
Capability register
"Current" means selected for its bounded role. It does not mean autonomous, generally capable, or safe outside that role.
| Project | Model | Base | Role | Evaluation | System evidence | Decision |
|---|---|---|---|---|---|---|
| Loading model evidence. | ||||||
Evaluation protocol
The lab follows the measurement principles of NIST AI RMF and HELM: context-specific scenarios, multiple metrics, repeatable settings, visible limitations, and independent human review. It is not certified by NIST, Stanford, MLCommons, or ISO.
| Axis | What it measures | Current signal | Promotion rule |
|---|---|---|---|
| Capability | Correct task completion on a versioned, role-specific suite. | Case pass rate and delta | Must meet the project threshold on untouched cases. |
| Boundary integrity | Authority, safety, privacy, and fail-closed behavior. | Gate pass rate | Hard gate. One critical boundary failure blocks promotion. |
| Reliability | Repeat stability, calibration, false positives, and recovery behavior. | Reported when repeated runs exist | Must remain within the approved variance across repeat runs. |
| Efficiency | Latency, model size, hardware, and operational admission. | Seconds, gigabytes, hardware | Must fit the target workstation and response budget. |
| Human review quality | Correction distance, evidence usefulness, and reviewer agreement. | Correction distance when measured | Human review remains required even when the measured score improves. |
No universal intelligence number. Scores compare versions only within the same task, suite version, settings, and hardware context.
Missing is not zero. An absent measurement is shown as not measured and cannot satisfy a promotion gate.
Aggregate scores cannot override gates. High capability never compensates for an authority or safety violation.
Reference methods: NIST AI RMF Measure, Stanford HELM, VHELM, HEIM, and ISO/IEC 25010:2023.
Measured growth
Capability and boundary scores are shown separately. Missing measurements remain unmeasured, not zero.
Loading benchmark history.
Public capability lab
Curated demonstrations show how implementation quality changes across releases. They are not private evaluator cases or hidden-test outputs.
Living implementation / v1.0.0-rc.2
An original falling-block puzzle turns the CodeForge contract into something visitors can operate, inspect, and verify. Here, a specimen means a reference build with exposed evidence. The internal BlockForge record separates model assistance, deterministic tests, and human publication authority.
Approved reference / v0.1.0-core
A compact collision laboratory exposes the fixed brick field, reflection rules, explicit game states, and publication boundary. Power-ups remain deliberately deferred so the approved core stays legible.
Human review build / v0.2.0-review
A fixed single-screen platformer makes movement, collision, enemy, scoring, and lifecycle requirements observable. The same seed reproduces the chamber, while pause-safe timing and a digest-bound input receipt make completed runs reviewable.
Candidate vertical slice / v0.1.0-core
A deterministic tower-defense board exposes placement, targeting, wave, economy, pause, win, and loss contracts. The first slice keeps one tower and one threat class so the strategic core can be reviewed before feature breadth.
Publication boundary: prompts, judges, reference implementations, and raw candidates from sealed evaluations remain private. Public reference builds are purpose-built for explanation.
Decision record
Every retained, rejected, or testing decision carries a reason and an evidence class.
Loading promotion decisions.
Promotion method
Capability is promoted one bounded role at a time. A model never inherits authority from a good demo.
Name the task, inputs, outputs, and forbidden actions.
Use deterministic tests, hostile probes, and untouched holdouts.
Retain evidence without changing source, hardware, or production state.
Measure useful completion, missed requirements, regressions, and unnecessary change.
Advance only when the full gate passes and the authority boundary stays intact.
Operating boundary
PlantForge cannot water a plant. CodeForge cannot commit, publish, or deploy. ImageForge cannot select or publish its own output. MeshForge cannot approve or export a generated asset. VisionForge cannot declare compliance or approve publication. These are structural constraints, not promises in a prompt.
PlantOS validates evidence and owns care-state decisions.
No actuation routeDisposable worktrees and authenticated evidence contain model output.
Human commit requiredVersioned workflows and frozen briefs produce candidates for review.
Human selection requiredLocal image-to-mesh systems produce bounded geometry candidates.
Human export requiredMultimodal models produce structured visual findings for review.
Human judgment requiredEvaluation history
Loading the evidence timeline.