The NLT Labs Pipeline
One sentence in.
Demo site deployed end-to-end.
Phase 1 researches and evaluates (six analysts in parallel; a Devil's Advocate, Number Auditor, and Buyer Simulator adversarial pass; then a PE Firm verdict). Phase 2 builds and deploys the demo, gated by four parallel QA judges (visual, engineering, content, post-deploy) feeding a deterministic Quality Reviewer. A FUND verdict (pe:fund) auto-triggers the build; founder retrospective + GitHub review keep humans in the loop before deploy.
What the pipeline now outputs
The pipeline is no longer abstract. Three demos live on nltlabs.ai today — all featured here.
Why This Exists
The machine behind every pipeline run.
Most companies spend months and $50,000 figuring out if an idea is worth building. NLT Labs runs the same evaluation through a nine-section rubric, with forced Optimist, Skeptic, and Pragmatist lenses, then a Devil's Advocate stress-test. A FUND verdict (typically ~6+ average after the full rubric) queues the build crew for a deployed demo site with evaluation notes attached.
The key innovation is research-first. Every agent builds on real competitive data, real customer quotes from Reddit, real market evidence. Nothing is invented cold. The Market Researcher cites primary sources. The Competitive Analyst pulls actual Crunchbase funding data. The Creative Director runs in parallel and emits the brand identity — name, tagline, positioning — grounded in the same evidence base. By the time the PE Firm scores the idea, it's working with substance, not temperature.
Hardware products (when explicitly in scope, default pipeline skews software-first) get a dedicated Product Designer that thinks like an industrial designer: use scenarios, mechanism specs, electrical schematics, real BOM pricing at 1K and 10K unit volumes. All before code or 3D. The 3D render prompt is derived from the mechanism spec, not invented.
The output is a demo deployed as a two-door site: a polished consumer experience for browsers, and a single-scroll evaluation-notes page (/inside) with real market data, unit economics, roadmap, capital ask, and a contact button. Hardware POCs add a rotating 3D model, lifestyle hero render, studio render, and a downloadable Tech Specs sheet. Four parallel QA judges (visual, engineering, content, post-deploy live-URL audit) gate the ship, and a deterministic Quality Reviewer aggregates them into one verdict — any blocked verdict blocks the ship, no LLM override. Everything grounded in what the agents actually found.
Meet the Team
29 specialists wired into workflows/*.yaml. Each one does one job.
Software-only evaluations skip the hardware crew (Product Designer, Product Renders, 3D Modeler, Engineering QA); the Zone C prototype lane only runs when an idea earns a demand signal. Dig Deeper is a follow-on loop, not a full-time seat on every idea.
Runs in parallel with the five analysts and owns brand identity: product name, tagline, positioning. The PE Firm reads this alongside the research briefs so the verdict has a name to fund — not just a thesis.
Finds the real market, with sources. No invented TAM numbers. Real customer voice from Reddit and reviews. Real industry data with citations.
Maps the competitive landscape with real funding data. Crunchbase raises in the space, what they built, why they raised. Grounds the PE Firm's valuation reality check.
Runs scope:deep sequential mode, not a shallow scan. Answers "can we actually build this?" The real questions: what tech, what team, what timeline, and what might blow up in year two.
Runs the numbers with primary sources required. Customer cost to acquire, lifetime value, break-even point. Broken math shows up in the PE verdict. Not as a rubber stamp.
Hunts for legal landmines with citations. Licensing requirements, data privacy exposure, liability. Primary source required. No guessing.
Before the PE write-up: attacks the thesis. Fact-checks quantitative claims, stress-tests financials, hunts hidden competitors, and rebuts each brief with evidence. The PE Firm reads this pass first, then owns the verdict.
The math gate. Inspects every quantitative claim across the briefs: TAM/SAM/SOM math, unit economics, contradictions between sections, missing formulas, implausible projections. Owns a broken-math hard cap that floors the PE Firm score when arithmetic is unsound. The reason "derived from cited numbers" actually means something.
An independent buyer, scored separately from the committee. Models whether a real customer would actually pay at the modeled price — and refuses by default: a would-buy verdict requires named evidence answering every objection. Its refusal caps the PE Firm score, so "they'll pay" has to be earned, not assumed.
The judge. Runs a nine-section Gate 1 rubric (forced Optimist / Skeptic / Pragmatist lenses), then lands on a verdict — invest, pass, or monitor — and summarizes the headline scorecard most people see: Market, Execution, Moat, and Timing, each with explicit justification.
Dig Deeper (conditional): when the PE verdict asks for more research, a targeted follow-up run answers specific questions, sometimes with a human gate, then the issue is re-scored. It does not replace the Devil's Advocate pass on a full evaluation.
Names the product, researches competitors, and writes a brief grounded in real market data. Saves a competitive-research.md with competitor table, pain points, category visual language, and differentiation angle. Every downstream agent reads it.
Hardware only. The industrial design brain. Runs 6 stages: use scenario narrative, form language rationale, mechanism spec, electrical schematics, sensory design, competitive contrast. Outputs an interactive BOM at 1K and 10K unit volumes, a three-size SKU strategy (S/M/L dimensions matched to the use case, retail and gross margin per tier), and a category-appropriate materials + regulatory cert sheet (FCC, FDA, UL — scoped to the product). The render and 3D prompts are derived from this spec. Not invented.
Anchors pricing to comparables before founder intuition. Pulls the BOM cost from Product Designer (hardware) or feature complexity, blends three signals (60% comparable retail, 25% BOM bounds, 15% modal proxy), and emits S/M/L tier prices. The /pricing page renders against this — not whatever the operator typed at intake.
Hardware only. Generates four photorealistic product renders — studio, lifestyle hero, exploded internals, and S/M/L product family — via Nano Banana (Gemini 2.5 Flash Image). An anchor-then-variant pattern locks subject, product, and scene identity across renders so the lifestyle hero matches the studio shot matches the exploded view. Containerized agent; on a failed Visual QA verdict the workflow re-runs it with the rejection rationale as negative-prompt context, up to 3 iterations. The rotating 3D model is a separate agent (next).
Hardware only. Derives an interactive, rotating .glb from the isolated studio angles via Meshy multi-image-to-3D, then a single-round fidelity-vs-source judge blocks a model that does not match the renders. Best-effort by design — a 3D failure never fails the ship. Powers the rotating GLB viewer on the Tech Specs page.
Reads competitive research before touching any layout. A Visual Differentiation section is mandatory: "competitors do X, we do Y, because Z." Every layout decision is grounded in something real.
Opens Zone C. Sets the pre-declared go/no-go demand thresholds for the idea — interested-over-exposed target, minimum exposure, the hard-commitment bar — and resolves whether a playable prototype is earned. The verdict math itself stays deterministic Python.
Zone C. Specs the core interaction loop and a typed mock-API contract for the prototype builder, so the clickable demo exercises the real job-to-be-done rather than a static mock. Self-skips unless the demand gate clears.
Zone C. Builds a front-end-only playable prototype (Vite + React + MSW with seeded fixtures) mounted at /app, so a visitor can actually click through the core loop. An independent Prototype QA verifies the clickable loop and re-runs up to 3 times; the whole lane self-skips unless the demand gate clears.
Research-first. Every line grounded in real customer pain points from the competitive-research.md. Hardware copy must match the actual mechanism. Reads the Market Pricing tiers so /pricing and the business-model panel use the same numbers everywhere.
Writes the 18-section GTM playbook: ICP, buyer persona, trigger events, lead sources, first-100 strategy, cold-call/cold-email/LinkedIn/partnership scripts, discovery questions, objection handling, offer/pricing/landing tests, design-partner criteria, 30/60/90 kill criteria. An independent GTM QA judge scores it against an anti-slop rubric and re-runs it until ≥16 of 18 sections are specific. Mounted at /gtm on the deployed site.
Independent vision judge. Scores every render against the DSG / Soft-TIFA rubric (anatomy, lighting, scene coherence, subject identity, brand-fit). Emits per-render verdicts; hard rejects feed back into the Product Renders regen loop. Cross-family by design — Claude Sonnet vision judging renders a different model family (Gemini) produced.
Independent engineering judge (hardware). Checks the built spec against the constraint spec — BOM completeness, power budget, thermal envelope, dimensional fit — and feeds a pass/fail into the ship verdict alongside the vision and content gates. Fast and cheap on Haiku.
Independent sales-readiness judge. Scores every customer-facing and investor page against the 100-pt operator rubric: word-count band, CTA count, proof anchors, comparison element, FAQ, banned-phrase scan, citations, quant artifacts, FATE sub-scorecard. Landing pages need ≥80, investor pages ≥85, or the ship blocks.
Builds the two-door site (consumer-facing + the 9-route /inside investor brief, including AI & NLT Platform), wires the Product Renders and the rotating 3D model plus the BomTable into the Product & BOM tab, then deploys under a custom subdomain and notifies Bill with the live link + a 1–5 star quality rating prompt. Also ships Path to Product and Tech Specs.
Operator-grade audit against the live URL. Greps for BUILDER: stubs, fetches every internal route for HTTP codes, validates citations, checks brand palette, scans for hallucinations. Local lint pass + live audit fail = ship blocked. Catches the class of bug where the dev preview is clean but the deployed site is broken.
Deterministic aggregator (no LLM). Reads Visual QA, Engineering QA, Content QA, Market Pricing, Builder, and Post-Deploy Audit verdicts and applies one rule: any blocked ⇒ blocked; ≥1 needs_revision ⇒ needs_revision; else complete. The subdomain-claim gate consults this signal before reverting a tentative claim.
The Output
Not a prototype. A demo site with full evaluation notes attached.
Every FUND-verdict idea is built into all of this. The /inside evaluation-notes page alone would cost $10,000+ from a consultant.
Self-Improvement
The system improves itself.
Every deploy feeds back into the pipeline. Over time, the system gets better at predicting which ideas produce good outcomes.
After each deploy, the Provisioning Agent prompts Bill for a 1–5 star quality rating on Telegram. Site quality, brief quality, anything off.
7 days later: a check-in. 🔥 Strong interest / 👍 Some / 😐 None / 🗑 Kill. Real market response from real people who saw the site.
Hot signals trigger a structured debrief. What worked? What resonated? That becomes the v2 brief, or a funded product.
All ratings and signals save to shared memory. Every agent recalls them at the start of each run. The pipeline learns what works, and what doesn't.
The feedback loop is what separates a pipeline from a machine that learns. Each rated POC makes the Creative Director's seed enrichment more calibrated, the PE Firm's scoring more predictive, and the whole system more likely to surface the ideas that actually become products.
Why It Matters
The old way costs a fortune. This doesn't.
Most ideas die not because they're bad, but because validating them is expensive. We changed that math.
"We're not replacing human judgment. We're running it at a scale and speed no human team could match. So the ideas that deserve to exist get a real shot."
Each AI agent has a single job and does it the way a specialist consultant would. The PE Firm scores through nine forced-lens sections, then lands on a verdict; the Devil's Advocate tries to tear the thesis apart with evidence. The Product Designer outputs a real BOM before anyone generates a 3D model. The Copywriter reads customer pain quotes before writing a headline. The result is an honest evaluation and a real demo deployment. Not a pitch deck.
Stack Claude Opus 4.8 + Sonnet 5 + Haiku 4.5 · Gemini 2.5 (Nano Banana + Pro judge) · Meshy image-to-3D · AgentForge orchestrator · tapps-brain shared memory · firecrawl + exa for research · GitHub App · Render + GoDaddy for deploy
Full colophon ↗Got an idea worth running through this?
One sentence in. The same pipeline that built the demos above evaluates it, then builds and deploys a demo if it earns a FUND verdict.