AI Tools Lab

Tutorial · Automation

Build an evidence-first AI content research pipeline.

Published September 8, 2026 · Original architecture · No affiliate links on this page

Automation is useful when it removes repetitive collection and formatting. It becomes dangerous when it automates the part that needs judgment: whether a product is actually good, whether a claim is supported, and whether a recommendation deserves a reader's trust.

Design rule: automate signals, records and drafts; keep product testing, material claims and early publishing behind a human gate.

The pipeline

public demand signals
→ opportunity records
→ scoring
→ evidence pack
→ content hypothesis
→ draft
→ factual/claim QA
→ human approval
→ publish
→ clicks/conversions/usefulness signals
→ learn and reprioritize

This architecture is intentionally boring. The database and evidence pack matter more than the text generator because they preserve what the system knew when a decision was made.

1. Collect demand, not just trends

A spike in mentions is not necessarily buyer intent. Keep separate signal types:

The best content opportunities often combine buyer intent with a product you can genuinely test.

2. Store every candidate as a record

A minimal opportunity table might contain:

id
discovered_at
product_a
product_b
intent
source_urls[]
problem_statement
target_reader
affiliate_available
evidence_status
score
state
notes

Use SQLite, Postgres or even a well-structured spreadsheet at the beginning. The important part is persistence: do not let promising ideas disappear inside chat history.

3. Score for revenue and editorial value separately

One score creates perverse incentives. Keep at least two:

A high commission should never convert a weak product into a positive recommendation. Commercial score determines whether a useful piece can also monetize—not whether it is useful.

4. Build an evidence pack before drafting

The drafting model should not browse randomly and improvise. Give it a compact evidence pack:

This makes hallucination easier to detect because the permitted factual surface is defined.

5. Generate a hypothesis, not “an article”

A good content hypothesis has a reader, decision and proof:

For: technical solo creators
Decision: ElevenLabs or local TTS
Proof: same 10 scripts + cost per accepted minute
Format: comparison article + 30s short demo
Success: qualified outbound clicks + strong completion/saves

Now the content exists to answer a decision, not to fill a publishing calendar.

6. Create a claim ledger

For every material statement, classify it:

Claim typeRequired evidence
Current price/featureOfficial vendor page, dated
Measured performanceReproducible direct test
User sentimentMultiple recent community sources; label as anecdotal
Commercial termsCurrent affiliate/network terms
“Best” / recommendationTransparent decision criteria and alternatives

If a claim cannot be placed in the ledger, either verify it or remove it.

7. Render multiple formats from one evidence object

The same verified record can generate:

This is the safe way to “scale content”: scale the reuse of verified evidence, not the number of unsupported claims.

8. Human QA should have a checklist

9. Measure the funnel, not vanity reach

Views are an attention metric, not a business result. Track the chain:

impression/view
→ retained attention
→ profile/page visit
→ outbound affiliate click
→ merchant conversion
→ revenue
→ refund/churn where relevant

For owned content, useful secondary signals include search impressions, scroll depth, return visits and email opt-ins. Revenue should ultimately reconcile to the affiliate network dashboard.

10. Make the system able to kill ideas

Automation that only produces more is incomplete. Add rules for stopping weak experiments: repeated poor retention across different hooks, high clicks with no conversions, product quality deterioration, bad economics, or excessive production effort.

A system that kills bad ideas quickly creates more value than one that generates 100 drafts per day.