The NLT Labs Pipeline

One sentence in.
Demo site deployed end-to-end.

Phase 1 researches and evaluates (six analysts in parallel; a Devil's Advocate, Number Auditor, and Buyer Simulator adversarial pass; then a PE Firm verdict). Phase 2 builds and deploys the demo, gated by four parallel QA judges (visual, engineering, content, post-deploy) feeding a deterministic Quality Reviewer. A FUND verdict (pe:fund) auto-triggers the build; founder retrospective + GitHub review keep humans in the loop before deploy.

29
Specialists · across both phases
2
Phases · 33-node DAG
~30–50 min
Wall-clock · evaluation + ship
$17–22
Hardware POC total · incl. QA judges
~6+
Typical FUND threshold (avg)

What the pipeline now outputs

The pipeline is no longer abstract. Three demos live on nltlabs.ai today — all featured here.

Open pipeline output ↗
Collage of demo sites deployed by the NLT Labs pipeline
Pawvlov demo site preview
Pawvlov

Hardware device that detects storms and rewards calm behavior so dogs unlearn thunder anxiety.

BidFast demo site preview
BidFast

Voice notes turned into branded, legally-compliant estimates — win the job before leaving the driveway.

PitchPad demo site preview
PitchPad

A driveway walk-through turned into a signed proposal before the truck pulls away — field-first for owner-operators.

Why This Exists

The machine behind every pipeline run.

01

Most companies spend months and $50,000 figuring out if an idea is worth building. NLT Labs runs the same evaluation through a nine-section rubric, with forced Optimist, Skeptic, and Pragmatist lenses, then a Devil's Advocate stress-test. A FUND verdict (typically ~6+ average after the full rubric) queues the build crew for a deployed demo site with evaluation notes attached.

02

The key innovation is research-first. Every agent builds on real competitive data, real customer quotes from Reddit, real market evidence. Nothing is invented cold. The Market Researcher cites primary sources. The Competitive Analyst pulls actual Crunchbase funding data. The Creative Director runs in parallel and emits the brand identity — name, tagline, positioning — grounded in the same evidence base. By the time the PE Firm scores the idea, it's working with substance, not temperature.

03

Hardware products (when explicitly in scope, default pipeline skews software-first) get a dedicated Product Designer that thinks like an industrial designer: use scenarios, mechanism specs, electrical schematics, real BOM pricing at 1K and 10K unit volumes. All before code or 3D. The 3D render prompt is derived from the mechanism spec, not invented.

04

The output is a demo deployed as a two-door site: a polished consumer experience for browsers, and a single-scroll evaluation-notes page (/inside) with real market data, unit economics, roadmap, capital ask, and a contact button. Hardware POCs add a rotating 3D model, lifestyle hero render, studio render, and a downloadable Tech Specs sheet. Four parallel QA judges (visual, engineering, content, post-deploy live-URL audit) gate the ship, and a deterministic Quality Reviewer aggregates them into one verdict — any blocked verdict blocks the ship, no LLM override. Everything grounded in what the agents actually found.

End-to-End Flow

The machine that runs from idea to URL.

Each box is a real specialist agent. Six analysts — market, competition, feasibility, finance, regulation, and brand — run in parallel against the idea; then three adversaries pile on — the Devil's Advocate stress-tests the bull case, the Number Auditor inspects every number for broken math, and the Buyer Simulator independently models whether a real buyer would pay; the PE Firm reads all three and issues the verdict. Dig Deeper only runs when the verdict asks for more research. Not on every idea.

Phase 1: Evaluation
~13–15 min wall-clock · 6 parallel analysts → Devil's Advocate + Number Auditor + Buyer Simulator → PE Firm
💡 The Idea 🎨 Brand · Creative name · tagline · positioning 📊 Market Size Primary sources · Reddit voice 🥊 Competition Crunchbase funding data ⚙️ Build Cost scope:deep sequential 💰 The Money Primary source required ⚖️ Legal Risk Primary source required 6 RUNNING IN PARALLEL SONNET · ADVERSARY 😈 Devil's Advocate fact-check · rebuts each brief SONNET · MATH GATE 🧮 Number Auditor broken-math hard cap SONNET · BUYER 🛒 Buyer Simulator would-buy · refuses by default OPUS · JUDGE 🏦 PE Firm 9 lenses · M/E/Mo/T summary VERDICT M:x E:x Mo:x T:x FUND verdict → POC queue (auto) ~6+ avg typical Not FUND → pass · dig deeper
pe:fund verdict → poc-ship auto-triggers · GitHub review as needed
Phase 2: Build
~17–30 min wall-clock · 18-agent build crew · 4 QA gates + regen loops · deterministic ship-readiness verdict
Software path
Hardware-only path
OPUS · DIRECTOR 🎯 POC Director brief · slug · downstream specs SONNET · HW ONLY 🏭 Product Designer BOM · mechanism · schematics ↻ MAX 3 REGENS 💵 Market Pricing comparables · BOM bounds 🖼️ Product Renders Nano Banana · 4 renders 🧊 3D Modeler Meshy · rotating .glb 🎨 Web Designer brand tokens · diff section ✍️ Copywriter research-first · pain-led 📣 GTM Operator 18-section playbook 6 IN PARALLEL · HW BOXES SKIP ON SOFTWARE SONNET · QA GATE 👁️ Visual QA DSG · per-render verdict HAIKU · QA GATE 🔧 Engineering QA constraint-spec checks (HW) SONNET · QA GATE 📋 Content QA 100-pt sales-readiness OPUS · CODER 🔨 Frontend Builder two-door site · /inside SONNET · QA GATE 🔍 Post-Deploy Audit live-URL · stubs · citations DETERMINISTIC · VERDICT ✅ Quality Reviewer aggregator · no LLM 🚀 Demo Site *.nltlabs.ai ZONE C · EARNED BY A DEMAND SIGNAL — SELF-SKIPS OTHERWISE 📡 Signal Planner demand-gate thresholds 🧩 Prototype Designer core-loop spec · mock API 🕹️ Prototype Builder Vite + React + MSW · /app ↻ Prototype QA + gateway verify · max 3 regens
Quality Gates · ship-readiness verdict

Four independent judges run in parallel after the build crew — vision, engineering, content, and live-URL — then a deterministic aggregator decides whether the demo ships. No LLM in the aggregator; the verdict is reproducible.

👁️
Visual QA
DSG / Soft-TIFA

anatomy · lighting · scene coherence · subject identity across renders

🔧
Engineering QA
Constraint spec

BOM · power · thermal · dimensions checked against the hardware constraint spec

📋
Content QA
100-pt rubric

CTAs · proof anchors · banned phrases · citations · FATE sub-scorecard

🔍
Post-Deploy Audit
Operator audit

live-URL routes · BUILDER: stubs · palette · hallucination scan

Quality Reviewer · deterministic aggregation

Any blocked verdict ⇒ ship blocked. ≥1 needs_revision ⇒ needs_revision. Otherwise complete. The subdomain-claim gate consults this signal before reverting a tentative claim.

Meet the Team

29 specialists wired into workflows/*.yaml. Each one does one job.

Software-only evaluations skip the hardware crew (Product Designer, Product Renders, 3D Modeler, Engineering QA); the Zone C prototype lane only runs when an idea earns a demand signal. Dig Deeper is a follow-on loop, not a full-time seat on every idea.

11 Phase 1 · 18 Phase 2 · synced 2026-08-21
P1
Phase 1: Evaluation 11 roles · ~13–15 min wall-clock
🎨
Creative Director
Claude Sonnet

Runs in parallel with the five analysts and owns brand identity: product name, tagline, positioning. The PE Firm reads this alongside the research briefs so the verdict has a name to fund — not just a thesis.

📊
Market Researcher
Claude Sonnet

Finds the real market, with sources. No invented TAM numbers. Real customer voice from Reddit and reviews. Real industry data with citations.

🥊
Competitive Analyst
Claude Sonnet

Maps the competitive landscape with real funding data. Crunchbase raises in the space, what they built, why they raised. Grounds the PE Firm's valuation reality check.

⚙️
Feasibility Analyst
Claude Sonnet

Runs scope:deep sequential mode, not a shallow scan. Answers "can we actually build this?" The real questions: what tech, what team, what timeline, and what might blow up in year two.

💰
Financial Modeler
Claude Sonnet

Runs the numbers with primary sources required. Customer cost to acquire, lifetime value, break-even point. Broken math shows up in the PE verdict. Not as a rubber stamp.

⚖️
Regulatory Analyst
Claude Sonnet

Hunts for legal landmines with citations. Licensing requirements, data privacy exposure, liability. Primary source required. No guessing.

😈
Devil's Advocate
Claude Sonnet

Before the PE write-up: attacks the thesis. Fact-checks quantitative claims, stress-tests financials, hunts hidden competitors, and rebuts each brief with evidence. The PE Firm reads this pass first, then owns the verdict.

🧮
Number Auditor
Claude Sonnet

The math gate. Inspects every quantitative claim across the briefs: TAM/SAM/SOM math, unit economics, contradictions between sections, missing formulas, implausible projections. Owns a broken-math hard cap that floors the PE Firm score when arithmetic is unsound. The reason "derived from cited numbers" actually means something.

🛒
Buyer Simulator
Claude Sonnet

An independent buyer, scored separately from the committee. Models whether a real customer would actually pay at the modeled price — and refuses by default: a would-buy verdict requires named evidence answering every objection. Its refusal caps the PE Firm score, so "they'll pay" has to be earned, not assumed.

🏦
The PE Firm
Claude Opus

The judge. Runs a nine-section Gate 1 rubric (forced Optimist / Skeptic / Pragmatist lenses), then lands on a verdict — invest, pass, or monitor — and summarizes the headline scorecard most people see: Market, Execution, Moat, and Timing, each with explicit justification.

Dig Deeper (conditional): when the PE verdict asks for more research, a targeted follow-up run answers specific questions, sometimes with a human gate, then the issue is re-scored. It does not replace the Devil's Advocate pass on a full evaluation.

P2
Phase 2: Build 18 nodes · build crew + QA gates + deterministic reviewer · ~17–30 min wall-clock
🎯
POC Director
Claude Opus

Names the product, researches competitors, and writes a brief grounded in real market data. Saves a competitive-research.md with competitor table, pain points, category visual language, and differentiation angle. Every downstream agent reads it.

🏭
Product Designer
Hardware
Claude Sonnet

Hardware only. The industrial design brain. Runs 6 stages: use scenario narrative, form language rationale, mechanism spec, electrical schematics, sensory design, competitive contrast. Outputs an interactive BOM at 1K and 10K unit volumes, a three-size SKU strategy (S/M/L dimensions matched to the use case, retail and gross margin per tier), and a category-appropriate materials + regulatory cert sheet (FCC, FDA, UL — scoped to the product). The render and 3D prompts are derived from this spec. Not invented.

💵
Market Pricing Researcher
New
Claude Sonnet

Anchors pricing to comparables before founder intuition. Pulls the BOM cost from Product Designer (hardware) or feature complexity, blends three signals (60% comparable retail, 25% BOM bounds, 15% modal proxy), and emits S/M/L tier prices. The /pricing page renders against this — not whatever the operator typed at intake.

🖼️
Product Renders
Hardware
Claude Opus

Hardware only. Generates four photorealistic product renders — studio, lifestyle hero, exploded internals, and S/M/L product family — via Nano Banana (Gemini 2.5 Flash Image). An anchor-then-variant pattern locks subject, product, and scene identity across renders so the lifestyle hero matches the studio shot matches the exploded view. Containerized agent; on a failed Visual QA verdict the workflow re-runs it with the rejection rationale as negative-prompt context, up to 3 iterations. The rotating 3D model is a separate agent (next).

🧊
3D Modeler
Hardware
Claude Opus

Hardware only. Derives an interactive, rotating .glb from the isolated studio angles via Meshy multi-image-to-3D, then a single-round fidelity-vs-source judge blocks a model that does not match the renders. Best-effort by design — a 3D failure never fails the ship. Powers the rotating GLB viewer on the Tech Specs page.

🎨
Web Designer
Claude Sonnet

Reads competitive research before touching any layout. A Visual Differentiation section is mandatory: "competitors do X, we do Y, because Z." Every layout decision is grounded in something real.

📡
Signal Planner
Zone C
Claude Sonnet

Opens Zone C. Sets the pre-declared go/no-go demand thresholds for the idea — interested-over-exposed target, minimum exposure, the hard-commitment bar — and resolves whether a playable prototype is earned. The verdict math itself stays deterministic Python.

🧩
Prototype Designer
Zone C
Claude Sonnet

Zone C. Specs the core interaction loop and a typed mock-API contract for the prototype builder, so the clickable demo exercises the real job-to-be-done rather than a static mock. Self-skips unless the demand gate clears.

🕹️
Prototype Builder
Zone C
Claude Sonnet

Zone C. Builds a front-end-only playable prototype (Vite + React + MSW with seeded fixtures) mounted at /app, so a visitor can actually click through the core loop. An independent Prototype QA verifies the clickable loop and re-runs up to 3 times; the whole lane self-skips unless the demand gate clears.

✍️
Copywriter
Claude Sonnet

Research-first. Every line grounded in real customer pain points from the competitive-research.md. Hardware copy must match the actual mechanism. Reads the Market Pricing tiers so /pricing and the business-model panel use the same numbers everywhere.

📣
GTM Operator
New
Claude Sonnet

Writes the 18-section GTM playbook: ICP, buyer persona, trigger events, lead sources, first-100 strategy, cold-call/cold-email/LinkedIn/partnership scripts, discovery questions, objection handling, offer/pricing/landing tests, design-partner criteria, 30/60/90 kill criteria. An independent GTM QA judge scores it against an anti-slop rubric and re-runs it until ≥16 of 18 sections are specific. Mounted at /gtm on the deployed site.

👁️
Visual QA
QA Gate
Claude Sonnet

Independent vision judge. Scores every render against the DSG / Soft-TIFA rubric (anatomy, lighting, scene coherence, subject identity, brand-fit). Emits per-render verdicts; hard rejects feed back into the Product Renders regen loop. Cross-family by design — Claude Sonnet vision judging renders a different model family (Gemini) produced.

🔧
Engineering QA
QA Gate
Claude Haiku

Independent engineering judge (hardware). Checks the built spec against the constraint spec — BOM completeness, power budget, thermal envelope, dimensional fit — and feeds a pass/fail into the ship verdict alongside the vision and content gates. Fast and cheap on Haiku.

📋
Content QA
QA Gate
Claude Sonnet

Independent sales-readiness judge. Scores every customer-facing and investor page against the 100-pt operator rubric: word-count band, CTA count, proof anchors, comparison element, FAQ, banned-phrase scan, citations, quant artifacts, FATE sub-scorecard. Landing pages need ≥80, investor pages ≥85, or the ship blocks.

🔨
Frontend Builder
Claude Opus

Builds the two-door site (consumer-facing + the 9-route /inside investor brief, including AI & NLT Platform), wires the Product Renders and the rotating 3D model plus the BomTable into the Product & BOM tab, then deploys under a custom subdomain and notifies Bill with the live link + a 1–5 star quality rating prompt. Also ships Path to Product and Tech Specs.

🔍
Post-Deploy Audit
QA Gate
Claude Sonnet

Operator-grade audit against the live URL. Greps for BUILDER: stubs, fetches every internal route for HTTP codes, validates citations, checks brand palette, scans for hallucinations. Local lint pass + live audit fail = ship blocked. Catches the class of bug where the dev preview is clean but the deployed site is broken.

Quality Reviewer
Verdict
Deterministic

Deterministic aggregator (no LLM). Reads Visual QA, Engineering QA, Content QA, Market Pricing, Builder, and Post-Deploy Audit verdicts and applies one rule: any blocked ⇒ blocked; ≥1 needs_revision ⇒ needs_revision; else complete. The subdomain-claim gate consults this signal before reverting a tentative claim.

The Output

Not a prototype. A demo site with full evaluation notes attached.

Every FUND-verdict idea is built into all of this. The /inside evaluation-notes page alone would cost $10,000+ from a consultant.

🏠
Consumer Homepage
Emotional, product-focused. Hero, pain statement, feature callouts, lifestyle imagery, and a CTA that converts.
⚙️
Product Experience Pages
How It Works, Dashboard, and Pricing. Not a feature list. A story about what changes for the customer.
🚪
/inside Evaluation Notes
New
Tabbed investor brief, 9 routes: Overview → Market Research → Competitive → Product & BOM → Financial Model → Engineering → Path to Product → Process → AI & NLT Platform.
🛣️
Path to Product
New
Derived roadmap with real timelines and costs. Capital ask from actual component pricing and development estimates. Not templated.
🔩
Tech Specs + 3D Model
Hardware
Hardware only. Four photorealistic renders (studio, lifestyle hero, exploded internals, S/M/L product family), interactive BomTable at 10K volume, three-size SKU strategy with breed-matched dimensions and per-tier margin, materials + cert sheet, dedicated Engineering tab (architecture, integrations, data model, security posture, scalability), and a rotating GLB model viewer.
📋
Business Plan
Full PE analysis, competitive table with funding data, headline scorecard (Market/Execution/Moat/Timing from the nine-lens rubric), and GTM. All designed in.
📱
Fully Responsive
Right on a phone, a tablet, and a widescreen monitor. No exceptions. Mobile-first from the start.

Self-Improvement

The system improves itself.

Every deploy feeds back into the pipeline. Over time, the system gets better at predicting which ideas produce good outcomes.

01
Quality Rating

After each deploy, the Provisioning Agent prompts Bill for a 1–5 star quality rating on Telegram. Site quality, brief quality, anything off.

📡
02
Market Signal Check

7 days later: a check-in. 🔥 Strong interest / 👍 Some / 😐 None / 🗑 Kill. Real market response from real people who saw the site.

📝
03
Structured Debrief

Hot signals trigger a structured debrief. What worked? What resonated? That becomes the v2 brief, or a funded product.

🧠
04
Hive Memory Update

All ratings and signals save to shared memory. Every agent recalls them at the start of each run. The pipeline learns what works, and what doesn't.

The feedback loop is what separates a pipeline from a machine that learns. Each rated POC makes the Creative Director's seed enrichment more calibrated, the PE Firm's scoring more predictive, and the whole system more likely to surface the ideas that actually become products.

Why It Matters

The old way costs a fortune. This doesn't.

Most ideas die not because they're bad, but because validating them is expensive. We changed that math.

Traditional approach
Strategy consultant $15,000–$40,000
Market research firm $8,000–$25,000
Design agency (MVP) $20,000–$60,000
Development team $30,000–$100,000+
Industrial designer $5,000–$20,000
Timeline 3–9 months
Total $78K–$245K+
NLT Labs pipeline
PE evaluation: 6 analysts + Devil + Auditor + Buyer Sim + PE Firm $3–6 in AI tokens
Hardware POC: build crew + Nano Banana renders + Meshy 3D $8–15 in AI tokens
Software POC: build crew $3–5 in AI tokens
QA gates: Visual + Engineering + Content + Post-Deploy judges $1.50–2.50 in AI tokens
Hosting (Render SSR + static CDN) Marginal · bundled in ops
Custom subdomain + deploy Automated under NLT umbrella
Timeline ~30–50 min wall-clock
Hardware POC total ~$17–22

"We're not replacing human judgment. We're running it at a scale and speed no human team could match. So the ideas that deserve to exist get a real shot."

Each AI agent has a single job and does it the way a specialist consultant would. The PE Firm scores through nine forced-lens sections, then lands on a verdict; the Devil's Advocate tries to tear the thesis apart with evidence. The Product Designer outputs a real BOM before anyone generates a 3D model. The Copywriter reads customer pain quotes before writing a headline. The result is an honest evaluation and a real demo deployment. Not a pitch deck.

Stack Claude Opus 4.8 + Sonnet 5 + Haiku 4.5 · Gemini 2.5 (Nano Banana + Pro judge) · Meshy image-to-3D · AgentForge orchestrator · tapps-brain shared memory · firecrawl + exa for research · GitHub App · Render + GoDaddy for deploy

Full colophon ↗

Got an idea worth running through this?

One sentence in. The same pipeline that built the demos above evaluates it, then builds and deploys a demo if it earns a FUND verdict.

Submit your idea → Browse portfolio