Living whitepaper / Updated August 30, 2026

ImageForge

A local visual-generation lab that turns art direction into repeatable still and motion workflows, then keeps model claims behind measured evidence and human approval.

Current releaseMotion v0.1
HardwareRTX 3080 Ti / 12 GB
Publication gateHuman approval required
01

Product truth

The output is not the system.

ImageForge is the workflow around the model: versioned direction, fixed controls, local execution, evidence capture, review criteria, and an explicit decision about whether an artifact can ship.

Local first
Weights and generated media remain on controlled hardware until a person selects an artifact for publication.
Comparable
Seeds, resolution, frame count, sampling controls, runtime, and outputs are recorded for each benchmark.
Reviewable
Identity, geometry, text, motion, artifacts, and accessibility are inspected before promotion.
02

Motion evidence

Two controlled Wan 2.2 benchmarks

Both clips use the same local model, seed, output size, frame count, frame rate, step count, CFG, sampler, and scheduler. The changed variable is source conditioning.

A / Text to video

Coffee motion study

A five-second generation testing object stability, liquid behavior, camera restraint, and commercial composition without a source image.

Render
134.7 s
Output
832 × 480
Frames
121 / 24 fps
Observed VRAM
10.7 GB peak

Finding: Stable cup, saucer, liquid stream, and composition. Requested steam was weak or absent.

B / Image to video

CourseForge preservation study

A source-conditioned test of face identity, cafe geometry, espresso-machine structure, CourseForge chrome, navigation placement, and text preservation.

Render
141.2 s
Output
832 × 480
Frames
121 / 24 fps
Observed VRAM
11.1 GB peak

Finding: Identity and major interface geometry remain coherent. Small text softens and some letterforms drift, so deterministic UI animation remains the production path.

03

AudioForge voice layer

Four local narration avatars

These samples use the local speech pipeline to turn frame text into reviewable narration. Each voice is a controlled avatar for tone testing, not a final learner-facing persona.

A / American male

Briefing Officer

Direct, concise, and procedural. Good for safety notes, task setup, and military-adjacent courseware.

B / American male

Range Instructor

More assertive and field-instruction flavored. Useful for checkpoints, warnings, and hands-on procedural demos.

C / American female

Calm Compliance

Steady and neutral. Good for policy, accessibility, compliance, and high-clarity enterprise training.

D / American female

Course Narrator

Warmer and more instructional. Better for walkthroughs, introductions, and longer lesson narration.

Regional coverage

Available now: American English and British English voices, plus Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin voice sets. Australian English has local candidate samples under review. African English accents are still research-only.

US / Female

American English

Comparable CourseForge narration sample using Kokoro voice af_heart.

US / Male

American English

Comparable CourseForge narration sample using Kokoro voice am_michael.

UK / Female

British English

Comparable CourseForge narration sample using Kokoro voice bf_emma.

UK / Male

British English

Comparable CourseForge narration sample using Kokoro voice bm_daniel.

AU / Candidate

Australian English

Ruby is the strongest current Orpheus Australian-English candidate. Human review is still required before it becomes a supported CourseForge narration voice.

ruby featured sample, 45.9 sec generation for 6.3 sec audio Chatterbox reference sample, 20.0 sec generation for 4.4 sec audio delta, 30.6 sec generation flynn, 26.5 sec generation ruby, 24.7 sec generation mason, 24.6 sec generation
04

Controlled comparison

What stayed fixed

ModelSeedStepsCFGSamplerScheduler
Wan 2.2 TI2V 5B20260830205uni_pcsimple

These are local capability probes on one workstation, not general performance guarantees. Video has no generated audio. Human temporal review remains required.

05

Current status

Capability ledger

The ledger records the latest verified state. Promotion requires evidence, not model availability.

Loading current status…

06

Progress record

What changed

  1. Loading progress history…
07

Next controlled cycle

Preserve the subject. Protect the text.

The next image-to-video cycle will keep model-generated motion away from UI labels and composite deterministic text back over the clip. Success means the motion can feel alive while the interface remains crisp, readable, and governed by source copy.