PromptZone - Leading AI Community for Prompt Engineering and AI Enthusiasts

Content briefs and outlines

A content brief fails in one of two ways: it is so thin the writer guesses ("write about X, 1500 words, include the keyword"), or so bloated it is a disguised first draft. Both produce content that reads like everyone else's — which, in a SERP that already has ten of those, is a page that never ranks.

These prompts build the middle thing: briefs and outlines derived from the actual search intent in front of you — the real SERP, the real keyword list, the real page that decayed. Every one asks the model to commit to decisions (which intent, which angle, which page owns which query) and to show the evidence for them, because a brief that hedges every choice just moves the guessing to the writer.

Curated

Assign one owner page per query

These pages from one site get impressions for overlapping queries. For each query group, name the page that should own it and the evidence — which page already performs better, which page's content actually satisfies the query, not which was published first. Then per page: what it keeps, what it cedes, and the change that enacts the decision — retitle away from the ceded query, add the section that justifies ownership, or consolidate and redirect if a page ends up owning nothing. Note where apparent cannibalization is actually fine because the queries differ more than their wording suggests. Pages with their query data: {{data}}

Fill in: Each page with the queries it gets impressions for, positions and clicks per query if available.

Known limits: GSC attributes queries to pages noisily — one-impression overlaps are noise, not cannibalization; weigh the volumes.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
PAGE A: /guides/email-subject-lines — "How to Write Email Subject
Lines That Get Opened". Queries: "email subject lines" (pos 18, 900 imps, 4
clicks), "how to write a subject line" (pos 9, 400 imps, 11 clicks), "subject
line best practices" (pos 12, 250 imps, 2 clicks)

PAGE B: /blog/50-email-subject-line-examples — "50 Email Subject Line Examples
by Industry". Queries: "email subject lines" (pos 21, 700 imps, 2 clicks),
"email subject line examples" (pos 6, 1,100 imps, 41 clicks), "subject line
examples for sales" (pos 8, 300 imps, 9 clicks)

PAGE C: /blog/email-open-rates — "What's a Good Email Open Rate? (Benchmarks)".
Queries: "average email open rate" (pos 7, 2,000 imps, 55 clicks), "email
subject lines" (pos 45, 60 imps, 0 clicks)

Model output (unedited)

Summary recommendation
- Page A (/guides/email-subject-lines) should own the how-to / best-practices intent and the broad query "email subject lines".
  Evidence: A ranks best for the general query (pos 18 vs B pos 21, C pos 45) and already performs well on the how-to queries ("how to write a subject line" pos 9, 400 imps, 11 clicks; "subject line best practices" pos 12).
- Page B (/blog/50-email-subject-line-examples) should own examples intent (all “examples” queries, incl. industry-specific).
  Evidence: B clearly owns example intent: "email subject line examples" pos 6, 1,100 imps, 41 clicks; "subject line examples for sales" pos 8, 300 imps, 9 clicks.
- Page C (/blog/email-open-rates) should own metrics/benchmark intent ("average email open rate").
  Evidence: C is the clear winner for that query (pos 7, 2,000 imps, 55 clicks) and its topical focus (benchmarks) matches user intent. C performs very poorly on subject-line queries (pos 45, 60 imps, 0 clicks), so it should not try to own that intent.

Why this split
- Intent differs: some users want instruction/best practices (A), some want concrete examples by industry (B), some want benchmarks/metrics (C). That difference is meaningful even when query wording overlaps — it's fine for multiple pages to appear for related but distinct intents as long as each page is optimized for a clearly different need.
- Empirical performance supports this: A leads on how-to/best-practice queries; B leads on examples queries; C leads on open-rate benchmarks.

Per-query ownership (short)
- "email subject lines" (ambiguous/broad) — owner: Page A. Evidence: A ranking and content match for general guidance; B serves example intent but ranks slightly worse; C irrelevant.
- "how to write a subject line" — owner: Page A (pos 9, 400 imps, 11 clicks).
- "subject line best practices" — owner: Page A (pos 12, 250 imps).
- "email subject line examples" — owner: Page B (pos 6, 1,100 imps, 41 clicks).
- "subject line examples for sales" — owner: Page B (pos 8, 300 imps, 9 clicks).
- "average email open rate" — owner: Page C (pos 7, 2,000 imps, 55 clicks).

Per-page actions (keeps, cedes, concrete change)

Page A — /guides/email-subject-lines
- Keeps: All how-to and best-practices intent (targets: "how to write a subject line", "subject line best practices", and the broad "email subject lines" general intent).
- Cedes: Deep examples-by-industry intent (ceded to Page B) and specific example search volume.
- Changes to enact decision:
  1. Make H1 and prominent H2s explicitly focused on strategy and best practices: e.g., H1 includes "How to Write Email Subject Lines" (already present — keep), H2s like "Best practices" and "Step-by-step checklist".
  2. Add a short "Quick examples" section (5–10 representative subject lines) to satisfy users who want a quick sample, but include a strong internal link and a sentence: "For a larger, industry-organized list, see 50 Email Subject Line Examples" (link to Page B). This prevents users wanting examples from bouncing and signals clear intent separation to search engines.
  3. Ensure meta title/description emphasize guidance/strategy (e.g., "How to Write Email Subject Lines — Tips, Best Practices & Checklist") so the page does not compete for pure example searches.

Page B — /blog/50-email-subject-line-examples
- Keeps: All example-driven intent, including industry-specific example queries ("email subject line examples", "subject line examples for sales").
- Cedes: Broad how-to / best-practices ownership (ceded to Page A).
- Changes to enact decision:
  1. Keep the current H1 and title focused on examples (ensure "Examples" is in H1 and meta title: e.g., "50 Email Subject Line Examples (by industry)").
  2. Add clear H2 anchors for each industry (including "Sales") to match long-tail example queries and capture featured snippets.
  3. Add a short "How to use these examples" intro paragraph that links to Page A for strategy/best-practices. Keep this short — avoid adding a long how-to section that would re-create A's content and reintroduce cannibalization.
  4. Optionally add structured data for lists/snippets if appropriate to improve appearance in SERPs.

Page C — /blog/email-open-rates
- Keeps: Benchmark and metric intent (owning "average email open rate" and related queries).
- Cedes: Any subject-line how-to or examples intent (ceded to Pages A and B).
- Changes to enact decision:
  1. Ensure H1/meta title remain focused on open rates (e.g., "What's a Good Email Open Rate? — Benchmarks & Benchmarks by Industry") and remove or de-emphasize any generic "subject line" optimization in titles or H1s.
  2. If there is a section that mentions subject lines, rework it to be explicitly about how subject lines impact open rates (data-driven), and link to Page A for best practices and Page B for examples. Example phrasing: "Subject lines strongly affect opens — for tips on writing them, see our guide; for example subject lines, see this post."
  3. Avoid duplicating A’s best-practices content or B’s examples; keep the page metric-focused.

Notes where cannibalization is acceptable
- The general query "email subject lines" is ambiguous — some searchers want examples, others want instruction. It’s acceptable (and often desirable) for both a guide (A) and an examples list (B) to rank for variants or for search to show both types of results. The key is to make each page's intent explicit in titles/H1s and to link between them so search engines and users see a clear separation of purpose.
- Small overlap (A includes a short examples block; B includes a brief how-to note) is fine and helpful for users — but avoid long duplicate sections that compete for the same query.

No pages need full consolidation/redirects
- All three pages have clear, non-empty ownership and measurable user interest. Do not consolidate or redirect any of them. Instead use internal linking, clearer headings/meta titles, and lightweight section adjustments to reduce harmful cannibalization and align each page to a distinct intent.

Priority checklist (what to implement first)
1. Page A: add short examples section + strong internal link to Page B; ensure best-practices H2s and meta title emphasize guidance. (High impact — clarifies ownership of the broad query.)
2. Page B: add H2 anchors for industries (including Sales); keep title explicitly "Examples"; add a short link to Page A. (High impact — secures example queries.)
3. Page C: remove any competing subject-line SEO signals; add a data-driven "subject lines & open rates" blurb linking to A/B. (Medium impact — stops confusion and supports both pages.)
4. Monitor rankings/clicks for the overlapping queries for 4–8 weeks and iterate — if one page continues to cannibalize another for a query the other should own, tighten phrasing and increase internal linking to signal clearer intent.

If you want, I can draft specific H1/H2 and meta title+description examples for each page and a short internal-link anchor text plan.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Brief a glossary term page that earns its URL

Brief a definition page for this term. The definition itself must be committed and quotable in the first sentences — but a page that stops there does not deserve a URL, so specify what makes this one referenceable: the distinctions from the terms it is confused with, the worked example that shows the term in use, and the boundaries — what the term does not cover, where practitioners disagree. List the related terms this page links to and the anchor context for each. State who links to definition pages in this field and what would make this one the one they cite. Term and field: {{term}}

Fill in: The term, the field it belongs to, and the neighboring terms it gets confused with.

Known limits: Glossary pages compete with dictionaries and Wikipedia — without the example and boundaries, this page loses to both.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-nano

Input fixture
TERM: churn rate

FIELD: SaaS metrics, for a site read by early-stage SaaS founders

CONFUSED WITH / NEIGHBORS: retention rate (the inverse, but time-window
subtleties differ), revenue churn vs customer churn (routinely conflated —
losing 2 customers of 100 is 2% customer churn but could be 30% revenue churn
if they were the big ones), gross vs net revenue churn (net can be negative
with expansion revenue), involuntary churn (failed payments vs decisions).
Practitioner disagreement: whether to annualize monthly churn by multiplying
by 12 (compounding says no).

Model output (unedited)

TERM: churn rate
FIELD: SaaS metrics (for early‑stage SaaS founders)

Definition (quotable, first sentence)
Churn rate is the share of your base you lose in a defined period, expressed as a percentage of the starting base. In SaaS, this can be measured as customers churned (customer churn) or revenue churn (revenue lost from existing customers), within a chosen time window.

What makes this one referenceable
- Distinctions from terms it is confused with
  - Retention rate vs churn rate: Retention is the complement of churn for the same base and window, but in practice definitions and windows differ, so they are not always exact inverses.
  - Revenue churn vs customer churn: Customer churn tracks the count of customers lost; revenue churn tracks the revenue lost from those customers. They can diverge when big customers leave or small customers churn in bulk.
  - Gross revenue churn vs net revenue churn: Gross churn ignores any revenue expansion; net churn accounts for expansion/upsell in the period. Net churn can be negative if expansions exceed losses.
  - Involuntary churn: This is churn caused by failed payments or payment issues, which some definitions exclude and others include depending on the measurement goal.

- Worked example (term in use)
  Scenario: One SaaS product with 100 customers at the start of the month; starting MRR is $10,000. During the month:
  - 6 customers cancel (customer churn = 6%).
  - The revenue lost from those cancellations equals 6 customers × $100 ARPU = $600.
  - Expansions/upsells add $1,200 in expansion revenue (expansion revenue).
  Calculation:
  - Gross customer churn rate: 6% (600/10,000) if expressed as revenue impact; but typically stated as customer count: 6/100 = 6%.
  - Gross revenue churn: $600 / $10,000 = 6.0%.
  - Net revenue churn: (600 − 1,200) / 10,000 = -$600 / $10,000 = -6.0% (net churn negative when expansions exceed losses).
  Takeaway: churn rate can be positive (losses exceed gains) or negative (expansion revenue outpaces losses).

- Boundaries (what the term does not cover)
  - It does not by itself measure new customer acquisition, activation, or overall growth; it only describes loss from an existing base within a period.
  - It is sensitive to the chosen base and window (monthly vs quarterly vs yearly; cohort vs flat base).
  - It does not prescribe how to treat involuntary churn; you must decide whether to include or exclude failed payments from the churn calculation.
  - It does not substitute for a full cohort or path-to-value analysis; it is one metric among many for understanding stickiness.

- Practitioner disagreement
  - Annualization approach: Should you annualize monthly churn by multiplying by 12 (simple but often inaccurate) or use a compounding method (e.g., 1 − (1 − monthly_churn)^12) which can give different results and implications for planning?

Related terms this page links to (anchor context)
- Retention rate (anchor: “See also: Retention rate”)
  - Context: complementary view to churn; discusses keeping customers within the same window and the inverse relationship.
- Customer churn rate (anchor: “See also: Customer churn rate”)
  - Context: specifically the percentage of customers lost, as opposed to revenue.
- Revenue churn rate (anchor: “See also: Revenue churn rate”)
  - Context: revenue-based measure; can be separately discussed as gross vs net.
- Gross revenue churn (anchor: “See also: Gross revenue churn”)
  - Context: revenue lost from churned customers ignoring expansions.
- Net revenue churn (anchor: “See also: Net revenue churn”)
  - Context: revenue after accounting for expansions/upsells.
- Involuntary churn (anchor: “See also: Involuntary churn”)
  - Context: churn due to payment failures; may be treated differently in analyses.
- Time window / annualization (anchor: “See also: Time window and annualization”)
  - Context: how the chosen period affects the churn figure and comparability.
- Churn vs. expansion / cohort analysis (anchor: “See also: Cohort analysis”)
  - Context: broader methods to understand churn dynamics over time and across segments.

Who links to definition pages in this field and what would make this one the one they cite
- Typical linkers in this field:
  - SaaS KPI glossaries and startup metric primers (founders and operators use as quick references)
  - Investor resources and pitch-deck glossaries (VCs and accelerators point founders to clean metric definitions)
  - SaaS analytics and dashboard tools’ documentation (for standard metric definitions in dashboards)
  - Founder communities and blog series on growth metrics (for practical interpretation and examples)
- What would make this page the one they cite:
  - A concise, quotable core definition that clearly distinguishes churn rate from similar terms
  - Clear, practical worked example showing both customer churn and revenue churn (including gross vs net)
  - Explicit boundaries and a concise note on where practitioners disagree (annualization)
  - Well-placed links to related terms with precise anchor texts and contexts
  - Up-to-date discussion of current practice, including involuntary churn and the difference between gross and net churn
  - A direct note about when to use monthly vs annual figures and how to interpret them in planning and forecasting

If you want, I can tailor this page to a specific tone (super concise vs. deeper) or adapt the worked example to your product’s typical ARPU and expansion patterns.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Build a content brief from a keyword and its SERP

Build a writer-ready brief for this keyword from the SERP evidence provided. Commit to: the dominant intent and the page type that serves it; the angle this page takes that the ranking pages do not, grounded in what the SERP actually lacks; the sub-questions the page must answer, ordered by how the searcher's need unfolds; and the evidence the writer must gather before drafting — the specific tests, numbers, or examples that make the page citable rather than another summary. Name what the brief deliberately excludes and why. Do not specify a word count; the coverage list is the length. Keyword: {{keyword}} SERP: {{serp}}

Fill in: The target keyword and the current SERP: ranking titles, what each page covers, and any visible SERP features.

Known limits: The angle is only as good as your SERP notes — a lazy SERP summary produces a brief for a SERP that does not exist.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
KEYWORD: home coffee roasting for beginners

SERP (top results, annotated):
1. "How to Roast Coffee at Home" - BeanScene — oven/popcorn-popper/dedicated-roaster
   methods overview, no cost breakdown, stock photos
2. "Home Coffee Roasting: A Complete Guide" - RoastLab — long, covers science of first/
   second crack well, gear list is 2023-dated, no green bean sourcing advice
3. "I Roasted Coffee in a Popcorn Popper for a Year" - r/Coffee thread — real experience,
   messy, top comment warns about smoke and chaff cleanup
4. Video: "Roasting Coffee in a Pan (Don't)" - 8 min, negative result content
5. "Best Home Coffee Roasters 2026" - GrindDaily — pure product roundup
Feature: People Also Ask — "Is roasting your own coffee cheaper?", "Is home roasting
safe?", "How long do green coffee beans last?"

Model output (unedited)

Brief for: “home coffee roasting for beginners”

1) Dominant intent and page type
- Dominant intent: Informational — “how to start roasting coffee at home” with strong transactional/utility elements (equipment buying guidance, starter roast profiles, and cost/safety tradeoffs).
- Page type that best serves this intent: Actionable beginner’s guide + starter plan — a single long-form how-to that combines: a plain-language decision checklist (method & budget), step-by-step tested roast recipes for 3 common methods, simple troubleshooting, sourcing & cost math, and safety/ventilation guidance. Think “teach-me-and-hand-me-a-plan-I-can-follow-for-my-first-10-roasts.”

2) Unique angle this page MUST take (what the SERP lacks)
SERP summary: existing pages give method overviews (BeanScene), deep science (RoastLab), anecdotal forum posts (Reddit), a “don’t do this” video (pan), and pure product roundups (GrindDaily). Missing from the top results:
- A practical, data-backed starter plan that combines:
  - clear startup-cost & per-pound economics (is it actually cheaper?)
  - actionable sourcing advice for green beans (where to buy, what to pick for beginners)
  - tested—repeatable—roast profiles (times/temps/visual cues) for the most common beginner methods (popcorn popper/air roaster, small drum/dedicated home roaster, stovetop pan — with a note to avoid pan roasting).
  - explicit safety and ventilation steps for apartment cooks (smoke mitigation, detector guidance)
  - a 10-roast practice plan with expected taste evolution and simple cupping checklist
None of the top results combine these practical, measurable elements into step-by-step, test-backed guidance for absolute beginners.

Positioning statement to lead copy: “A beginner’s starter kit you can follow on day one: decide a budget, buy beans, execute three tested roast profiles (air & drum), manage smoke safely, and evaluate results — plus a 10-roast practice plan with cost math so you know whether this will save money.”

3) Primary audience & tone
- Audience: complete beginners (no roasting experience), including apartment dwellers, home-brewers curious about flavor control, and budget-conscious hobbyists.
- Tone: plain, reassuring, evidence-focused. Prioritize checklists, step-by-step instructions, photos/diagrams, and numbered roast logs. Avoid heavy chemistry.

4) Coverage list — what the page must include (this is the length; do all)
A. Quick decision checklist (first screen / “Should I try this?”)
  - One-line goals: flavor exploration, cost savings, hobby.
  - Key constraints: ventilation, appliance rules in rentals, patience/time.
  - Clear recommendation paths by budget/space/goal (e.g., apartment + <$75 → air popper + venting vs. committed hobbyist → $200–400 dedicated roaster).

B. Safety & logistics up front
  - Ventilation checklist: windows + fan + DIY box + using a range hood; recommended minimum air changes or practical proxy (run kitchen fan + open window + box fan pointed out).
  - Smoke detector, carbon monoxide note, chaff fire risk and how to avoid it (no open flames under chaff piles; cool/dispose).
  - Apartment-specific precautions (notify landlord? check lease? use balcony where legal).
  - Short “Do not” list: pan roasting for beginners (link to video that fails), avoiding makeshift high-fire devices.

C. Starter budget & cost math (missing in SERP)
  - Upfront cost bands with representative SKUs:
    - Low budget: $20–80 (hot-air popcorn popper / basic air roaster)
    - Mid: $100–400 (entry dedicated home roaster, e.g., 150–250g batch machines)
    - Higher prosumer: $500+ (large-capacity drum roasters)
  - Per-pound economics worked example:
    - Inputs to show: green bean price range (collect low/median/high: e.g., commodity/commodity specialty), roast loss %, electricity use per roast (estimated watts × minutes), amortized gear cost per lb over 1 year at X roasts/week, time cost.
    - Produce 3 example scenarios (cheap home roasting vs. buying specialty roasted vs. buying commodity roasted) with total $/lb and break-even roast frequency.
  - Recommendation: when it makes financial sense (thresholds).

D. Where to buy green beans and what to buy (missing in SERP)
  - Reliable vendor types: specialty roasters selling green, green coffee brokers, local co-ops, online retailers (list examples to research).
  - What to look for: origin, processing method, crop year, moisture content/packing, roast-level intent (light/medium/dark), defect rates.
  - Starter bean picks: 3–5 accessible varieties that show clear roast-stage flavor changes (e.g., washed Ethiopian, washed Colombian, Brazil natural). Say why each is good for beginners.
  - Storage guidance for green beans: shelf life numbers (collect sources: typical 6–12 months depending on storage), recommended containers and temps.

E. Gear primer (short & practical — don’t duplicate long product roundups)
  - Gear comparison table to include: method, typical batch size, control level (time/temp), cleanup difficulty, smoke level, cost band, suggested models to test.
  - Minimum extras: scale (0.1g), thermometer or roast probe (if applicable), cooling tray or colander + fan, roasting logbook template, oven mitts, fire extinguisher.
  - What to buy first vs later.

F. Tested starter roast profiles (this is the most important evidence section)
  - For each of 3 methods (air-popcorn/air-roaster, small dedicated electric roaster, stovetop/drum — but include “don’t pan roast” caution), present 3 reproducible profiles: light, medium, dark.
  - For each profile provide:
    - Starting weight of green beans (g)
    - Equipment used (exact model)
    - Ambient conditions (room temp, altitude note if relevant)
    - Roast timeline: time to first crack, time to second crack (if reached), total roast time, and a target bean surface temperature if available (or probe reading).
    - Observable cues (color, smell, audible cracks, surface oil)
    - Weight loss % (pre/post roast)
    - Chaff volume/weight (if measured)
    - Expected flavor notes for that bean at that roast level
    - Photo set: green → end of roast (close-up)
  - Writer must run (or commission) at least 3 replicate roasts per method/profile to produce mean and range for times and % weight loss.
  - Provide an explicit “first-roast recipe” that a beginner can follow exactly (times/vent, cooling steps, when to stop).

G. Cooling, degassing, storage, and how long roasted beans last
  - Exact cooling methods and target time-to-cool.
  - Degassing timeline and how it affects brewing (peak flavor ~2–7 days depending on roast; cite tests).
  - Storage: airtight container + room temp vs fridge/freezer (practical recommendations and what’s backed by tests).
  - Freshness shelf-life numbers: peak window, decline curve (numbers to collect from studies or tests).

H. How to taste & evaluate (simple cupping & brewing tests)
  - Beginner cupping protocol (small-scale): dose, grind, brew, evaluation checklist (acidity, sweetness, body, aftertaste, off-flavors) and what to look for after light vs dark roasts.
  - A/B test plan to compare your roast to a shop-bought roast (same bean if possible): brew method, variables to hold constant, what differences to expect by roast level.
  - Example tasting notes for the sample beans and profiles above.

I. Troubleshooting & FAQ (answer SERP “People Also Ask” plus real-world problems)
  - Is roasting your own coffee cheaper? (answer with the scenario numbers and break-even)
  - Is home roasting safe? (concise safety recap + stats/risks)
  - How long do green coffee beans last? (sourced numbers)
  - Common issues: underdevelopment/souring, tipping/burnt edges, smoky/off-flavors, chaff clumping — symptoms + fixes.
  - Clean-up checklist (chaff & oil residues).

J. Progress plan: first 10 roasts and skill milestones
  - A 10-roast plan with specific goals per roast: observe cracking, control roast speed, reproduce a profile, tweak to adjust sweetness/acidity, record & compare.
  - What to expect at roasts 1–3 (learning the machine), 4–7 (dialing flavors), 8–10 (consistent repeatable roast).

K. Visuals & UX elements to include
  - Photo pairs for each profile (green vs finished), chaff pile photo, ventilation setup diagram (window+fan+filter), roast log screenshot, simple cost math table, “first roast recipe” printable checklist, short embedded video (30–90s) of a full roast on an air roaster.
  - Call-outs for safety rules and “stop here” warnings.

L. Recommended next steps & further learning
  - Where to learn roast theory (link to deeper science pages for readers who want more)
  - Community resources (forums, local roaster clubs)
  - When to upgrade gear and what to consider

5) Sub-questions the page must answer, ordered by how the searcher’s need unfolds
1. Should I even try roasting at home (benefits vs costs/risks)?
2. Is it safe and legal where I live / in my apartment?
3. What budget & method should I choose for my space and goals?
4. Where do I buy green beans and which beans should a beginner start with?
5. What exact equipment and accessories do I need day-one?
6. How do I roast a first batch — step-by-step recipe I can follow?
7. How do I immediately cool, store, and brew the roasted beans?
8. How will the flavor change over the next days — when is it best to brew?
9. How do I evaluate my roast and what troubleshooting steps should I take?
10. Is it cheaper than buying roasted coffee once I include all costs?
11. What are common mistakes and how to avoid them?
12. How do I progress beyond beginner level?

6) Evidence & tests the writer must gather before drafting (to make content citable)
- Real-world roast logs and photos:
  - Perform (or commission) at least 3 replicate roasts for each method/profile (air popcorn or entry air roaster; entry-level dedicated roaster; note: do not recommend pan unless documenting failure). Record time-to-first-crack and time-to-second-crack (if occurring). Provide mean +/- range.
  - Measure and report weight loss % for each roast (weigh green and roasted beans).
  - Photograph beans at start and immediate end-of-roast under consistent lighting + include color swatches or reference cards.
  - Record chaff weight or estimated volume after each roast.

- Equipment metrics:
  - For 3 representative machines (one in each cost band), measure typical batch size, roast time, ease of cleanup (time to clean), smoke output anecdote (qualitative), and energy use estimate (W × minutes) or cite manufacturer wattage and calculate energy per roast (kWh).
  - Note control features (temperature probe, airflow control, presets).

- Cost calculations:
  - Gather current price data for green beans at three price points (commodity, mid, specialty) and for roasted retail equivalents.
  - Calculate amortized equipment cost per lb for realistic usage patterns (e.g., 1 roast/week, 3 roasts/week, 7 roasts/week) over 1 year and 3 years.
  - Include electricity cost assumptions (local kWh $) and time estimate (labor).

- Freshness & degassing:
  - Measure / cite how flavor evolves post-roast: collect a small tasting panel (or at minimum a single consistent taster) with brews at day 0, day 2, day 7 for the same roast. Record perceived changes. If new testing is infeasible, cite published tests/studies or reputable roaster data.

- Safety evidence:
  - If possible, cite or run an indoor smoke/particulate proxy (or reference building science sources) for smoke levels produced by air popper vs electric home roaster vs open pan (or gather authoritative sources). At minimum, collect vendor recommendations and documented incidents if any.
  - Gather manufacturer safety specs and common-sense ventilation setups used by hobbyists (prefer peer-reviewed or government guidance if available).

- Sourcing & bean selection:
  - Collect purchase links and product pages for 4–6 reputable green bean vendors and sample product pages (origin, crop year, moisture).
  - Note price per lb and sample roast recommendations from vendors.

- Tasting & sensory: prepare 3 sample tasting notes (for the starter bean set) at the 3 roast levels so readers can compare expectations.

- External authoritative references to cite:
  - Specialty Coffee Association (SCA) guides on cupping/standards
  - Manufacturer pages for recommended roasters (wattage, safety)
  - Any research or reputable sources on roasted bean degassing and shelf-life
  - Consumer safety advice for indoor smoke/carbon-monoxide if relevant

7) Examples the writer should include (concrete, citable)
- “First-Roast Recipe”: Popcorn popper, 50 g green washed Colombian, expected first crack at 4–5 minutes, stop at 7:30 for light roast; weight loss ~15%; tasting notes: bright acidity, floral notes.
- “Midline Recipe”: Dedicated 150g roaster, 150 g green Brazilian, time to second crack at 9:30, stop at 10:30 for medium-dark; weight loss ~18%; tasting notes: chocolate, nuttiness.
- Cost scenario table: Example A — green beans at $6/lb, roasted store coffee at $14/lb, amortized roaster cost $200 at 2 roasts/week → break-even at X months (run the numbers).
- 10-roast practice plan with expected sensory checkpoints.

8) Visual & UX priorities for the writer/producer
- Lead with a succinct “Do this first” box (method by budget + safety checklist).
- Use side-by-side photos of roast stages and a printable one-page first roast checklist.
- Include an interactive calculator or downloadable spreadsheet for the cost math (optional but valuable).
- Provide downloadable roast log template.

9) What this brief deliberately excludes (and why)
- Deep chemical/academic explanations of roast chemistry at molecular level — excluded because beginners need practical roast control and sensory cues, not advanced theory. Link to RoastLab or SCA resources for readers who want deeper science.
- Exhaustive product roundup and affiliate-style “best roasters 2026” comparison — excluded because GrindDaily and other roundups already cover long lists; instead provide representative models by cost band and the decision criteria for choosing.
- Professional/commercial roasting workflows and industrial roaster calibration instructions — excluded because they require permits/ventilation and fall outside home beginner scope.
- Pan-roasting guides as a recommended method — exclude detailed "how to pan roast" instructions; include a clear caution and link to the “don’t” video for demonstration. Provide pan-roast as an option only to explain why it’s not beginner-friendly and what the risks/downsides are.

10) Metadata & SEO cues (practical)
- Primary keyword to target: home coffee roasting for beginners (use in title/H1, first paragraph, alt text for photos).
- Secondary keywords to surface: how to roast coffee at home, best green coffee beans for home roasting, popcorn popper coffee roast, is home roasting cheaper, home roaster safety.
- Suggested slug: /home-coffee-roasting-beginners
- Suggested H1: Home Coffee Roasting for Beginners: A Safe, Budget-Minded Starter Plan
- Suggested meta description (one line): Practical step-by-step guide for beginners — choose your method, buy green beans, run three tested roast recipes, manage smoke safely, and see if home roasting saves you money.

11) Sources & expert voices to get / cite
- Specialty Coffee Association resources (cupping, roast standards)
- 2–3 reputable home-roaster manufacturers (for wattage, batch, safety specs)
- 3 green-bean sellers (pricing and product pages)
- 1–2 experienced home roasters (interview for practical tips and apartment ventilation hacks)
- Optional: local building or fire safety guidance on indoor smoke/particulate limits (for legal/safety claims)

Final note to writer
- Make the guide actionable above all else: a complete beginner should be able to read the first section and confidently set up a first roast with the provided recipe and safety checklist. Back every prescriptive claim with a data point or test result (times, weight loss, cost numbers, photos). Where you cannot test (e.g., indoor particulate counts), clearly label as “reported anecdotally” and cite sources — do not make untested quantitative claims.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Cluster a keyword list into pages

Cluster these keywords into pages: each cluster is a set one page can rank for together because a searcher typing any of them wants the same thing. Name each cluster's page (working title and the primary keyword), list its member keywords, and state the shared need that justifies the grouping. Keywords that look similar but want different things go in different clusters — say which word told you. Mark keywords too weak or too far from the site's topic to deserve a page as skip, with the reason. Do not invent search volumes; if volumes are provided, use them to order clusters, not to create them. Keywords: {{keywords}}

Fill in: The raw keyword list, one per line, volumes optional.

Known limits: Wording-based clustering — two phrasings whose SERPs Google treats as one topic may still land in separate clusters; merge on evidence.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
KEYWORDS (from a site about backyard chickens):
how many chickens should i get
chicken coop size per chicken
how much space do chickens need
best chickens for beginners
friendliest chicken breeds
chicken breeds for eggs
how many eggs does a chicken lay per week
do chickens lay eggs in winter
why did my chickens stop laying
chicken stopped laying eggs molting
how cold is too cold for chickens
do chickens need a heat lamp
winterizing a chicken coop
can chickens eat tomatoes
what can chickens not eat
chicken safe table scraps list

Model output (unedited)

1) Working title: How many chickens should I get — planning flock size & space
Primary keyword: how many chickens should i get
Member keywords:
- how many chickens should i get
- chicken coop size per chicken
- how much space do chickens need
Shared need (why these belong together): Someone typing any of these is planning a backyard flock and needs a single page that helps them decide how many birds to keep, how much indoor/outdoor space they’ll need, and how to size/build a coop/run per bird. 
Which word signaled intent differences: "how many" signals a planning/quantity decision, while "coop size per chicken" and "how much space" signal the spatial requirements — together they answer the same planning question.

2) Working title: Best chicken breeds — friendly, beginner and egg-laying options
Primary keyword: best chickens for beginners
Member keywords:
- best chickens for beginners
- friendliest chicken breeds
- chicken breeds for eggs
Shared need: Searchers want help choosing which breed(s) to buy — evaluating temperament, ease-of-care for new keepers, and egg productivity — so one comparison/selection page that covers “best for beginners,” “friendliest,” and “best layers” fits.
Which word signaled intent differences: "friendliest" focuses on temperament; "chicken breeds for eggs" focuses on productivity. Both are breed-selection factors and belong on a single breed-choice page.

3) Working title: How many eggs will my chickens lay — rates, seasons, and troubleshooting
Primary keyword: how many eggs does a chicken lay per week
Member keywords:
- how many eggs does a chicken lay per week
- do chickens lay eggs in winter
- why did my chickens stop laying
- chicken stopped laying eggs molting
Shared need: These queries are about egg production: expected output, seasonal variation, and reasons/solutions when laying declines. A single egg-production page that explains normal rates, winter slowdown, molting, broodiness, disease, diet/stress causes, and remediation serves all these searches.
Which word signaled intent differences: "molting" indicates the specific cause someone suspects, while "do chickens lay eggs in winter" is about seasonal norms — both are troubleshooting/expectation aspects of egg production.

4) Working title: Keeping chickens warm — cold tolerance, heat lamps and winterizing
Primary keyword: how cold is too cold for chickens
Member keywords:
- how cold is too cold for chickens
- do chickens need a heat lamp
- winterizing a chicken coop
Shared need: These are winter-care questions: how cold is safe, whether supplemental heat is needed, and how to prepare a coop for cold weather. A practical how-to page on cold-weather care addresses all three.
Which word signaled intent differences: "heat lamp" shows someone considering active heating options; "winterizing" implies a broader preparedness checklist — both fall under cold-weather care.

5) Working title: What can chickens eat — safe foods and table scraps
Primary keyword: what can chickens not eat
Member keywords:
- what can chickens not eat
- can chickens eat tomatoes
- chicken safe table scraps list
Shared need: These queries are about feeding safety: which foods are safe or toxic and what scraps are acceptable. A single, clearly organized food-safety and safe-scraps page (with quick dos/don’ts and a list) satisfies this intent.
Which word signaled intent differences: "tomatoes" is a specific food inquiry; the others are general food-safety queries — all map to the same feeding-safety need.

Skip (no pages recommended)
- None. All provided keywords are on-topic and specific enough to justify being included in one of the pages above.

If you want, I can propose URLs and a suggested H2 structure for each page so content can be created/optimized to rank for the grouped keywords.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Critique a draft outline against its target query

Critique this outline against the query it targets. Walk the searcher's need through the outline in order: where does the outline answer it, what stands between the searcher and that answer, and what would make them hit back before reaching it. Flag sections that serve the writer instead of the searcher — history nobody asked for, product pitches ahead of the answer, padding sections that repeat a stronger one. Flag what the query implies that the outline never addresses. Verdict per section: keep, move (where to), cut, or missing (add what). Be specific enough that the writer can act without re-deriving your reasoning. Outline: {{outline}} Query: {{query}}

Fill in: The draft outline with a line on each section's content, and the target query.

Known limits: Critiques structure, not facts — a well-ordered outline of wrong information passes this review.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-nano

Input fixture
QUERY: is a heat pump water heater worth it

DRAFT OUTLINE:
1. The History of Water Heating — from kettles to modern tanks
2. What Is a Heat Pump Water Heater? — how it moves heat instead of making it
3. Our Company's Installation Process — how we quote and install
4. Benefits of Heat Pump Water Heaters — efficiency (3-4x), federal tax credit
5. Drawbacks — higher upfront cost ($1,800-3,200 installed vs $600-900), needs
   space and above-freezing location, slower recovery, cools the room it's in
6. Cost Comparison Over 10 Years — worked example: family of 4, electric
   resistance vs heat pump, payback in 3-4 years at average rates
7. Why Choose Us — testimonials
8. Conclusion — restates benefits

Model output (unedited)

Overall read: The outline covers the core facts about heat pump water heaters (HPWHs) and includes a useful cost-payback section, but it leans too heavily on writer-promotional sections, and it fails to deliver a clear, decision-focused answer to “is it worth it?” for the typical homebuyer. It also omits several practical decision factors (climate, usage patterns, electricity costs, installation constraints) and doesn’t provide a crisp verdict or scenario-based guidance. Below is a section-by-section verdict with actionable edits, followed by a recommended revised outline.

Section-by-section verdicts (keep / move / cut / missing)

1) The History of Water Heating — from kettles to modern tanks
- Verdict: Cut (or move to a very brief note in an intro)
- Why: The user query doesn’t require historical context. It adds length and invites filler rather than addressing whether HPWHs are worth it.
- Action: Remove or replace with a single sentence in the intro: “Here’s what makes HPWHs worth considering for modern homes.”

2) What Is a Heat Pump Water Heater? — how it moves heat instead of making it
- Verdict: Keep; Move up
- Why: Essential for readers who don’t know what a HPWH is. It should appear early so readers understand the product being assessed.
- Action: Place as section 2 (immediately after a short intro), or make it section 1. Keep concise, with a simple diagram or bullets.

3) Our Company's Installation Process — how we quote and install
- Verdict: Move (or cut)
- Why: Promotional content that creates friction for someone deciding whether HPWHs are worth it.
- Action: Move to a later section focused on “how to evaluate installers” or a generic, non-brand-specific installation considerations section. If retained, make it a short, non-promotional addendum at the end.

4) Benefits of Heat Pump Water Heaters — efficiency (3-4x), federal tax credit
- Verdict: Keep
- Why: Core positive case. Important to quantify benefits and incentives, but should be framed with caveats (works best in certain climates, usage patterns).
- Action: Keep, but add nuance: actual efficiency varies by climate, usage, and temperature settings; mention energy-factor vs. real-world performance.

5) Drawbacks — higher upfront cost ($1,800-3,200 installed vs $600-900), needs space and above-freezing location, slower recovery, cools the room it's in
- Verdict: Keep
- Why: Necessary counterpoint.Readers need to know when it may not be worth it.
- Action: Keep and enrich with regional/usage caveats (e.g., in very cold climates, in tight spaces, or in spaces that require heating).

6) Cost Comparison Over 10 Years — worked example: family of 4, electric resistance vs heat pump, payback in 3-4 years at average rates
- Verdict: Keep
- Why: This is the section most readers rely on to decide. It should be robust and transparent.
- Action: Expand with multiple scenarios (high vs. low hot-water usage, high vs. low electricity costs, with/without tax credits, with and without potential maintenance costs). Clearly label assumptions and provide a sensitivity range (e.g., payback 2–7 years depending on rates and usage).

7) Why Choose Us — testimonials
- Verdict: Move (or cut)
- Why: Promotional content that distracts from answering the query.
- Action: Move to the end as a “case studies / customer stories” section or remove entirely. If kept, present as third-party-style evidence (without heavy marketing language) and only after the decision framework is laid out.

8) Conclusion — restates benefits
- Verdict: Missing (rewrite)
- Why: The conclusion should deliver a direct verdict on the query and summarize decision criteria. As written, it restates benefits without answering “is it worth it?”
- Action: Replace with a decision-focused conclusion: a clear yes/no conditioned on factors, plus quick decision rules and a short set of next steps (e.g., factors to check before buying, questions to ask a seller).

Additional gaps the outline never addresses (missing but should be added)
- A clear decision framework: climate/region (cool vs warm), electricity rates, hot-water usage, tank size, and installation constraints. Provide a short checklist of “does this apply to you?”
- Real-world performance caveats: how climate and space heating affect performance; how much hot water you actually use; how to ensure the HPWH’s location doesn’t cause unwanted cooling or condensation.
- Sizing and installation basics: correct tank size, required clearances, electrical requirements, and the potential need for a dedicated circuit or load management.
- Maintenance and reliability: typical maintenance needs, warranty coverage, refrigerant handling, service availability.
- Alternatives and comparisons: how HPWH stacks up against other options (electric storage vs. tankless, gas-fired options, solar-assisted water heating) in different scenarios.
- Incentives and total cost of ownership: current federal/state incentives, local rebates, and how to claim them; lifetime operating costs vs. upfront costs.
- Practical guidance on choosing an HPWH: what specs to compare (EF/Uniform Energy Factor, first-hour rating, noise, space requirements), installer qualifications, and read-through of a typical service contract.
- Quick, decisive verdicts for common scenarios: e.g., “Worth it if you have electric heating with high rates in a warm climate and you use a lot of hot water; not cost-effective if you have cheap electricity and low hot-water use or you’re in a very cold climate without space to locate the unit.”

Recommended revised outline (actionable)

- Introduction: answer the question at a high level and set expectations
  - One-paragraph verdict with caveats
  - Quick outline of what affects worthwhileness (climate, usage, electricity price, installation space)

- 1) What is a Heat Pump Water Heater? (essential background)

- 2) When is it worth it? Decision framework
  - Climate/temperature considerations
  - Hot-water usage patterns
  - Electricity costs and incentives
  - Space, installation constraints, and noise
  - Sizing and efficiency expectations

- 3) Benefits (with caveats)
  - Efficiency ranges, real-world performance
  - Tax credits and incentives
  - Environmental impact considerations (optional)

- 4) Drawbacks and risks
  - Upfront cost, space, installation, climate sensitivity
  - Recovery time and impact on the space being cooled/heated
  - Maintenance and reliability notes

- 5) Cost comparisons and payback (with multiple scenarios)
  - Scenario A: typical family of 4 in a warm climate, high electricity cost
  - Scenario B: moderate usage, moderate electricity cost
  - Scenario C: low usage, cheap electricity, in a cool climate
  - Include payback ranges and break-even points; clearly state assumptions

- 6) How to choose and install (informational, not promotional)
  - Sizing, model features to compare
  - Installer qualifications and questions to ask
  - Installation considerations and potential pitfalls
  - Incentives and how to claim them

- 7) Real-world considerations and tips
  - Placement, space, ventilation, noise
  - Maintenance schedule and warranty tips
  - Common myths vs. facts

- 8) Conclusion
  - Clear verdict based on the decision framework
  - Quick next steps (e.g., if you want help sizing and comparing options, provide a checklist or questions to bring to a contractor)

- Optional appendix (non-promotional)
  - Customer stories or testimonials (if kept, ensure they’re clearly branded as independent evidence and not promotional)
  - Quick reference table of costs, payback ranges, and conditions

Key data points and language to include (so the writer doesn’t have to re-derive)
- Typical cost range: installed HPWH vs traditional electric storage; include current ranges and disclaimers about regional variations.
- Typical energy savings: “3–4x more efficient than electric resistance heating” with caveats.
- Payback range: roughly 2–7 years dependent on usage and electricity prices; provide ranges rather than a single number.
- Tax credits/incentives: summarize current federal tax credit and mention regional rebates.
- Climate guidance: note that HPWHs perform best in spaces that stay above about 40°F (4°C) and that very cold climates may reduce benefits unless installed with appropriate design considerations.

How to act on this guidance (concise checklist for the writer)
- Remove or relocate sections 1 and 7 (history and promotional company pitches) or replace them with a concise intro and a non-promotional installer-selection guide.
- Move section 2 to be immediately after the intro.
- Rework section 8 (Conclusion) to deliver a direct verdict with decision criteria, not just a restatement of benefits.
- Expand section 6 to include multiple scenarios and a transparent set of assumptions.
- Add a new decision framework section (climate, usage, electricity costs) and a sizing/installation guide.
- Keep essential benefits and drawbacks but present them with caveats and practical implications.
- Avoid product pitches; ensure every section serves the reader’s decision-making needs first.

If you’d like, I can provide a concrete, rewritten outline following this structure, with suggested copy for each section and a sample Q&A to help the writer respond to common reader questions.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Decide the page type for a mixed SERP

This keyword's SERP is split between page types, which means Google has not settled the intent — or serves more than one. From the evidence: name each intent present and which results serve it; judge which is winning (positions and features, not result counts); and commit to the page type we should build, with the argument. If the honest play is serving the secondary intent because the primary is unwinnable for us, say so explicitly. State what SERP change would flip your call. One decision, defended — not a hedge. Keyword: {{keyword}} SERP: {{serp}} Our situation: {{context}}

Fill in: The keyword, the annotated SERP, and your site's situation (what you sell or publish, existing authority).

Known limits: A mixed SERP is a moving target — the arbitration is valid for the snapshot you provide, and worth re-running quarterly.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
KEYWORD: dog dna test

SERP: #1 Embark product page; #2 "The 4 Best Dog DNA Tests, After 8 Months of
Testing" - WirecutterLike; #3 Wisdom Panel product page; #4 "Are Dog DNA Tests
Accurate? We Asked Geneticists" - VetJournal; #5 "Dog DNA Tests Compared: Embark vs
Wisdom Panel" - PupReview; #6-8: two more product pages and a Reddit thread "was
your dog DNA test worth it". Shopping carousel above organic; PAA: "How accurate
are dog DNA tests?", "Is Embark better than Wisdom Panel?", "Are cheap dog DNA
tests worth it?"

OUR SITUATION: independent dog-health content site; we do not sell tests; we have
a vet advisory board and an existing high-traffic guide on dog breed traits. We
can buy and run all four major tests on staff dogs.

Model output (unedited)

Short answer / decision
- Primary intent present (winning): Transactional — users looking to buy a dog DNA test. Evidence: shopping carousel above organic results and product pages in #1, #3, #6–8.
- Secondary intent (realistic target for us): Commercial investigation / “best-of” review / head-to-head comparison. Evidence: #2 (Wirecutter-like “4 Best…” at #2) and #5 (“Embark vs Wisdom Panel”) occupying high organic positions.
- Tertiary intent: Purely informational/scientific (accuracy, methodology). Evidence: #4 (“Are Dog DNA Tests Accurate?”) and PAA queries asking accuracy/comparisons/value; Reddit represents experiential/social proof.
- Our one decision: Build a data-driven, vet-reviewed commercial-investigation page — a Wirecutter-style comparative review + original head-to-head accuracy testing. Rationale below.

Why (short): Google’s SERP signals show buying intent is dominant (shopping carousel + product pages at top), but the highest non-product organic real estate is reserved for “best-of / review” type content — and that slot is winnable by an authoritative, in-depth comparison. We can’t out-compete manufacturers on their product pages or the shopping carousel, so the honest play is to own the commercial-investigation intent (the best-of/comparison slot) with original testing and credible vet expertise.

Mapping intents to results
- Transactional / Purchase:
  - #1 Embark product page
  - #3 Wisdom Panel product page
  - #6–8 other product pages
  - Shopping carousel (above organic)
- Commercial investigation / Review / Best-of:
  - #2 “The 4 Best Dog DNA Tests, After 8 Months of Testing” (Wirecutter-like)
  - #5 “Dog DNA Tests Compared: Embark vs Wisdom Panel” (PupReview)
- Informational / Scientific / Accuracy:
  - #4 “Are Dog DNA Tests Accurate? We Asked Geneticists” (VetJournal)
  - PAA questions: “How accurate are dog DNA tests?”, “Is Embark better…?”, “Are cheap dog DNA tests worth it?”
- Experiential / Social proof:
  - Reddit thread “was your dog DNA test worth it”

Which intent is winning (positions & features)
- Winner: Transactional (buy) — because the shopping carousel sits above organic results and product pages appear in the top organic slot and multiple positions. That’s a strong signal Google is prioritizing purchase fulfillment.
- Runner-up: Commercial-investigation / comparison — occupies the top non-product organic real estate (#2) and answers the PAA/comparison queries. This is the slot we can realistically win.

Why we should not target the transactional slot
- Manufacturer/product pages + shopping feature advantage (merchant schema, price feed, direct purchase CTA) are intrinsically hard for a content-only site to beat.
- We don’t sell tests — there’s no upside in trying to outrank product pages for “buy” queries when Google surfaces a shopping carousel.

What we should build (one decision)
- Build: A long-form, data-driven comparative review titled something like “The Best Dog DNA Tests (2026): Head‑to‑Head Accuracy Tests, Vet Analysis & Buying Guide.”
- Core differentiator: original head-to-head testing on our staff dogs plus vet/geneticist commentary and transparent methodology.
- Key elements to include:
  - Clear lead summary: best overall, best for health screening, best budget, best for breeders, etc.
  - Original test dataset: results from all four major tests run on the same sample dogs (purebreds, known mixes, intentionally blinded samples), with side‑by‑side breed detection, confidence scores, health variant reporting, turnaround time, cost, and privacy terms.
  - Methodology & limitations: exact sampling process, lab handling, sample size, statistical caveats (we can and should be honest if sample size is limited).
  - Expert analysis: vet advisory board commentary, independent geneticist quotes on variant detection and clinical utility.
  - Practical advice: when a test is useful, when it isn’t, how to interpret results, and recommended use-cases.
  - Comparison table and visual charts for quick scanning.
  - PAA-targeted FAQs that directly answer “How accurate…”, “Is Embark better…”, “Are cheap tests worth it?”
  - Strong E-E-A-T signals: author bios, vet board, methods appendix, raw data download, citations to peer-reviewed literature.
  - SEO & UX extras: schema for review, clear “buy vs research” CTAs (if linking to purchase), fast page speed, mobile-first layout.

Why this will win
- Matches the high-positioned “best-of” intent that Google already rewards (#2).
- Uses our concrete strengths: vet board, ability to buy and run tests, and existing traffic on breed traits.
- Provides what users and PAA want: accuracy information + comparison + recommendation — all in one canonical resource.
- Brands and merchants cannot easily replicate independent third-party testing and transparent methods, which helps E-E-A-T and linkability.

If the honest play is serving secondary intent, say so
- Explicit: Yes — the primary transactional intent is technically winning but essentially unwinnable for us (shopping carousel + manufacturer pages). The honest play is to serve the secondary commercial-investigation intent (comparisons/best-of) that is both prominent and winnable.

What SERP change would flip this call (i.e., when we would instead build a different page)
- Pivot to a pure informational/scientific page (deep-dive on accuracy and methodology) if Google begins to favor informational authority over transactional/comparison intent. Concrete flip triggers:
  - Shopping carousel removed or downgraded AND
  - Top organic positions (1–3) are dominated by authoritative informational articles (VetJournal-style pieces, research summaries) rather than product pages or “best-of” reviews — e.g., #1 becomes an authoritative “Are Dog DNA Tests Accurate?” piece with featured snippet and high PAA prominence around accuracy.
- In that scenario we would build a rigorously sourced, larger-sample accuracy study and meta-analysis — focusing on scientific evaluation over purchase advice.

Practical next steps (minimum viable page plan)
1. Run all four major kits on a small, intentional sample set (2–4 dogs: purebred with records, an obvious mix, and one edge-case) and document.
2. Draft a review structure: TL;DR recommendations, comparison table, deep-dive sections, FAQ (PAA-driven), methodology appendix.
3. Get vet and (if possible) independent geneticist quotes; prepare downloadable raw results and methods.
4. Launch as the definitive comparative review and promote to dog-owner communities, vets, and journalists to acquire links and social traction.

Bottom line: Don’t try to outrank product pages for “buy” intent. Build a high-E-E-A-T, original-data comparison/review that answers accuracy questions and buying decisions — that’s the slot Google has shown is accessible and the one our assets let us dominate.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Evidence checklist before writing a product review

Before this product review is written, list the evidence it needs to be worth publishing: the specific uses the reviewer must actually perform, the measurements or observations to record during them, the photos only real use can produce, and the comparison points against the named alternatives a reader will be deciding between. For each item, state what claim in the review it will support. Mark which items are non-negotiable — without them the review is a spec-sheet rewrite and should not be published. The checklist tests whether a review is possible, not how to phrase one. Product and context: {{product}}

Fill in: The product, who reads your reviews, and which alternatives readers ask about.

Known limits: A checklist cannot manufacture access — if nobody will actually use the product for weeks, no phrasing fixes that.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-nano

Input fixture
PRODUCT: Aeropress XL for a coffee gear review site whose readers are
home brewers deciding between it, the original Aeropress, and a French press.

CONTEXT: Reviews on this site historically rank when they include measured brew
data; readers in comments consistently ask about capacity ("finally enough for
two?"), whether existing filters/accessories fit, and cleanup time. The site owns
the original Aeropress and a Bodum French press. Review will be written by one
person over however long is needed.

Model output (unedited)

Below is a rigorous, evidence-focused checklist you can use before writing the Aeropress XL review. It’s designed to ensure the reviewer actually uses the product, gathers comparable data, and can credibly compare it to the original Aeropress and a Bodum French press. Each item lists the claim it supports, notes on what to record, and flags whether it’s non-negotiable for publication.

Categories and items

A. Specific uses the reviewer must perform (with the claims they support)

1) Maximum capacity test (two servings in one cycle)
- What to do: Fill the XL to its stated maximum capacity with coffee and water and press as you would for two equal servings; capture any overflow or leaks.
- Measurements/observations to record: total beverage yield (ml), any overflow incidents, number of presses needed to finish, cup-to-cup consistency.
- Photos to accompany: device filled to max line, the press in action, final two servings in cups.
- What claim it supports: The XL delivers two usable servings without overflow and with predictable yield.
- Non-negotiable: Yes.

2) Two-serving workflow and repeatability
- What to do: Reproduce two equal servings using the same coffee and recipe, ideally in same session or consecutive sessions.
- Measurements/observations: per-cup yield (ml), total time from start to pour, any variance between the two cups.
- Photos: side-by-side two cups, scale/measurements if used.
- Claim: Two consistent servings are feasible with repeatable results and similar extraction.
- Non-negotiable: Yes.

3) Single large cup feasibility test
- What to do: Brew one larger cup (e.g., 350–500 ml) using the XL without sacrificing extraction quality.
- Measurements/observations: total yield, brew time, grind setting used, dose, taste notes.
- Photos: single large cup in use and after pouring.
- Claim: The XL can produce a satisfactory single large cup; not just two small servings.
- Non-negotiable: Yes.

4) Filter options and performance test (paper vs metal or other filters)
- What to do: Brew using at least two different filters available for Aeropress (paper and metal if you have them) and compare flow, clear vs cloudy results, and taste.
- Measurements/observations: flow time (seconds per 50–100 ml), clogging/channeling incidents, clarity of brew, flavor differences.
- Photos: close-up of each filter in the basket during use; resulting cup.
- Claim: The XL supports multiple filter types and delivers noticeable differences (or similarities) in flow and flavor.
- Non-negotiable: Yes.

5) Accessory and filter compatibility test with existing gear
- What to do: Attempt to fit any filters/accessories you already own (original Aeropress filters, any available adapters) and document fit, seal, and any needed adjustments.
- Measurements/observations: fit quality, any wobble/leaks, ease of use, required tweaks.
- Photos: installed accessories in place, any alignment marks or gaps.
- Claim: Existing filters/accessories either fit cleanly or require adapters; compatibility is documented.
- Non-negotiable: Yes.

6) Cleaning and maintenance time/effort
- What to do: After each test, disassemble and clean as you would in normal use; record the time and steps.
- Measurements/observations: total cleanup time, number of parts, water/detergent usage, ease of rinsing, any stubborn residues.
- Photos: dirty parts before cleaning, clean parts reassembled.
- Claim: Cleaning effort/time is comparable to (or better/worse than) expectations for the device and its accessories.
- Non-negotiable: Yes.

7) Pressing action and mechanical feel
- What to do: Assess the feeling and effort needed to press, including any stiffness, texture, or wobble in the plunger mechanism.
- Measurements/observations: subjective ease-of-press (qualitative), any mechanical play, and whether pressure remains smooth as the chamber empties.
- Photos/Videos: optional but helpful video of pressing in action.
- Claim: The pressing action is smooth and controllable without excessive force or leakage.
- Non-negotiable: Yes.

8) Grind size range and extraction tolerance
- What to do: Test a few grind sizes (finer, medium, coarser) to see how robust the XL is across a range of extractions.
- Measurements/observations: dose consistency, extraction signs (body, clarity, sweetness), time; any channeling or under-extraction symptoms.
- Photos: grind setup and any visual cues in the bed.
- Claim: The XL tolerates a reasonable grind range with predictable results.
- Non-negotiable: Yes.

9) Temperature handling and stability (optional but valuable)
- What to do: Note the water temperature at brew start and how quickly the beverage cools to drinking temperature.
- Measurements/observations: starting temp (°C), temp after brewing (°C), temp after 5–10 minutes.
- Photos: thermometer readings or notes.
- Claim: Temperature stability supports consistent extraction; differences vs alternatives are documented.
- Non-negotiable: No (optional but recommended if you have the tools).

10) Taste and sensory comparison with the original Aeropress
- What to do: Brew the same coffee and recipe in the XL and in the original Aeropress under the same conditions; compare flavor, body, acidity, sweetness, aftertaste.
- Measurements/observations: sensory descriptors for each device; any perceived differences in mouthfeel and aroma.
- Photos: prepared cups side-by-side.
- Claim: The XL yields a distinct taste/texture profile compared with the original Aeropress.
- Non-negotiable: Yes.

11) Taste and sensory comparison with the Bodum French press
- What to do: Brew a comparable coffee with the Bodum French press using a method you can replicate, and compare with XL results.
- Measurements/observations: sensory descriptors; body and aftertaste comparison; ease of use vs cleanup.
- Photos: cup(s) in isolation and side-by-side with others.
- Claim: The XL offers a different profile or advantages/disadvantages relative to a French press in practice.
- Non-negotiable: Yes.

12) Build quality and durability notes
- What to do: Note materials, finish, threading, seals, and any signs of wear, while using the device over several sessions.
- Measurements/observations: fit, finish, cosmetic wear, any loosening or gaskets needing replacement.
- Photos: close-ups of threads, seals, and any wear areas.
- Claim: Build quality is robust/fragile as observed; maintenance implications are described.
- Non-negotiable: Yes.

13) Practical footprint and storage observations
- What to do: Assess the device’s footprint, weight, and how it fits in your typical home-bar setup (cabinet height, storage space, travel).
- Measurements/observations: dimensions, weight, how it fits with kettle, scale, and cups in your setup.
- Photos: device in typical storage/usage context.
- Claim: Real-world footprint and storage practicality are clear for home brewers with typical setups.
- Non-negotiable: No (optional but useful).

14) Photos checklist (in-use, setup, and comparison)
- What to do: Take a consistent set of photos showing setup, live use (with two servings and one large cup), cleaning steps, and side-by-side comparisons with the original Aeropress and French press.
- Measurements/observations: clear visuals that back the data (fill levels, cap, filter, and cup).
- Photos to include (at minimum): unboxed/assembled shot, filled/mounting shot, pressing action, final pours, cleaning/reassembly, and side-by-side comparison.
- Claim: Visual evidence strengthens all capacity, compatibility, and usage claims.
- Non-negotiable: Yes.

B. Comparison points against the named alternatives (for each point, note the claim it supports)

- Capacity and serving size
  - Data to collect: max usable volume, typical two-serving yield, single large-cup yield.
  - Claims supported: XL’s real-world capacity and serving options vs original Aeropress and Bodum French press.

- Filter compatibility and impact on taste/flow
  - Data: filter-fit, flow time, clarity, and flavor differences with each filter type.
  - Claims: Filtration options and their effect on taste and ease-of-use.

- Ease of use and pressing effort
  - Data: required force, smoothness, and any bleeding/leaks during press.
  - Claims: User-friendliness relative to the original Aeropress and French press.

- Cleanup effort
  - Data: cleaning time, number of pieces, residue handling.
  - Claims: Maintenance burden relative to alternatives.

- Flavor/taste profile
  - Data: sensory notes for each device with the same beans/recipe.
  - Claims: How the XL compares in taste structure (body, acidity, sweetness) to the other two.

- Build quality and durability
  - Data: observed materials, tolerances, wear signs.
  - Claims: Longevity expectations vs the other devices.

C. Non-negotiable items (items necessary for publishable evidence)

- Capacity and serving tests (maximum capacity and two-serving yield) — this prevents a spec-sheet rewrite.
- Two-serving workflow (repeatability and yields) — demonstrates real-world use.
- Single large-cup test — shows flexibility beyond two small servings.
- Filter compatibility test (paper vs metal) — directly addresses reader questions about filters.
- Accessory/filter compatibility with existing gear (original Aeropress, Bodum items) — reader-relevant compatibility data.
- Cleaning time and process — critical for maintenance considerations.
- Pressing action and mechanical feel — user experience essential for a device-centric review.
- Grind size range and extraction tolerance — supports measured brew data and practical guidance.
- Taste/sensory comparison vs the two named alternatives (original Aeropress and French press) — core decision-making data for readers.
- Photos demonstrating setup, use, capacity, and cleaning — visual evidence to accompany claims.
- Build quality and durability notes — informs long-term value.
- Optional but strongly recommended: a minimal set of temperature-related observations if you have the tools — adds depth to extraction discussion.

Notes and tips for execution
- Use the same coffee and roast for all device comparisons to control variables.
- When possible, run each test with a repeatable recipe (dose, water mass, and temperature kept consistent).
- If you can, include a brief blind or blinded tasting note for the taste comparisons to minimize expectation bias.
- Document any anomalies (e.g., unexpected leaks, clogging, or unusually slow flow) and explain them in the review.
- Keep a simple data log: date, coffee, dose (g), water (g), brew time, yield (ml), temperature (°C), and a one-line sensory note per test.

If you cover all the above items, you’ll have a publishable, evidence-backed Aeropress XL review that clearly informs home brewers choosing between the XL, the original Aeropress, and a Bodum French press. If any of the “non-negotiable” items are missing, the review risks reading like a spec sheet rather than a practical, test-based assessment.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Find an angle a saturated SERP has not taken

The SERP for this keyword is saturated with competent pages. Generate five genuinely different angles for a new page — different in what the reader walks away with, not in framing or title phrasing. For each: what it requires that we have and competitors did not use (data, access, testing capacity), which searcher it serves better than every ranking page, and its honest weakness. Then pick one, argue it against the strongest alternative, and state what evidence must exist before writing starts. If no angle clears the bar with our actual resources, say the honest answer is not to write this page. Keyword, SERP, our resources: {{context}}

Fill in: The keyword, what currently ranks, and what your site uniquely has — data, expertise, testing ability.

Known limits: Differentiation costs execution — an angle you cannot resource is a worse outcome than not writing the page.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
KEYWORD: how to negotiate rent

SERP: five competent listicle/guide pages (major personal-finance sites) all
covering: research comparables, time your ask to lease renewal, offer longer
lease for lower rent, get it in writing. One NYT piece behind paywall. Reddit
thread at #7 with real scripts people used. No page shows outcome data.

OUR RESOURCES: we run a renters' rights newsletter with 40k subscribers we can
survey; a staff member negotiated rent successfully 3 times and kept the email
threads; we have a lawyer on retainer who reviews content; no proprietary
market-rate data; medium site authority in the renting niche.

Model output (unedited)

Quick summary: You can win this SERP only by offering something the other guides don’t — empirical outcomes, landlord perspective, legally enforceable language, or a reproducible tested toolkit. Given your assets (40k newsletter, one staffer’s email threads, lawyer on retainer), five distinct, viable page angles follow. Each is different in what the reader actually walks away with (not just a new headline).

1) Evidence-first: “Which rent-negotiation tactics actually work — data from real renters”
- What the reader walks away with: quantified success rates for specific tactics (e.g., “asking at renewal,” “offering a longer lease,” “sending an email script X”), broken down by context (city size, landlord type), plus guidance on expected realistic outcomes.
- What we must have that competitors didn’t use: a structured survey of our 40k subscribers (and possibly follow-up interviews), plus the staffer’s real negotiation threads to validate responses.
- Which searcher it serves best: evidence-seeking renters who want to know “what’s most likely to work for me” and want numbers, not guesswork.
- Honest weakness: self-selection and self-report bias; outcomes aren’t causal (people who use tactic X may differ); limited geographic representativeness.

2) Tested script playbook: “Scripts that got rent reduced — A/B-tested templates (email + phone)”
- What the reader walks away with: copy-and-paste scripts tailored to scenarios (renewal, hardship, comps) that are shown to produce higher reply/concession rates; an interactive script-builder that adapts tone and asks.
- What we must have that competitors didn’t use: live testing (sending versions to landlords or orchestrating a field trial) or at least a large corpus of real scripts + outcomes from subscribers/staffer; analytics showing reply/concession rates per script.
- Which searcher it serves best: renters who want a ready-to-send, high-probability message and don't want to “wing it.”
- Honest weakness: ethically and legally sensitive to run field tests without consent from landlords; small sample tests may not generalize; landlord variability huge.

3) Legally enforceable concessions guide: “How to get enforceable rent concessions and avoid landlord backtracking”
- What the reader walks away with: precise wording and clauses to use (addenda), examples of binding agreements, how to document, and what to do if a landlord reneges — localized guidance per major jurisdiction.
- What we must have that competitors didn’t use: lawyer-reviewed, jurisdiction-specific templates and checklists, plus the staffer’s exchange as a real example of enforceable communication.
- Which searcher it serves best: renters who secured a promise and want to lock it in, or who fear landlord will renege — people who care about enforceability.
- Honest weakness: high complexity and need for localization; may require many regional variants and risks over-legalizing something many landlords won’t sign.

4) Behavioral negotiation playbook: “Psychology-first tactics landlords respond to”
- What the reader walks away with: a set of persuasion techniques (reciprocity, framing, anchoring, scarcity) tailored to landlords with scripts and roleplay examples — i.e., how to structure the ask to increase empathy and reduce resistance.
- What we must have that competitors didn’t use: input from negotiation/behavioral-science experts, ideally small experiments or at least qualitative landlord interviews showing differential responses to framing.
- Which searcher it serves best: renters confident in conversational negotiation who want to improve persuasion skills rather than just copy a script.
- Honest weakness: mostly theory unless backed by outcome data; performance varies widely by landlord personality and market pressure.

5) Landlord-business case toolkit: “How to pitch a rent cut so landlords see it as a profit-preserving decision”
- What the reader walks away with: a package (short tenant business brief, cost/benefit numbers, vacancy-cost calculator, optional small rent-deferral agreement) they can present to landlords to make the economic case for keeping them at a lower rent.
- What we must have that competitors didn’t use: interviews/survey data from landlords/property managers about their financial thresholds and decision triggers; templates that show projected savings vs. vacancy cost.
- Which searcher it serves best: higher-rent tenants in buildings where landlords are small owners/operators and will respond to numbers; tenants aiming to negotiate with property managers or owner-operators.
- Honest weakness: access to landlords is required to produce credible templates; big landlords/management companies may not care about a tenant’s pitch.

Choosing one to build: I recommend the Evidence-first angle (#1).

Why this over the strongest alternative (Tested script playbook #2)
- Strength of alternative: Scripts have immediate utility — people can copy-paste and act now — and Reddit threads already show real scripts, so a tested-script piece could rank well.
- Why evidence-first is better for us:
  - Unique asset fit: You can tap your 40k newsletter to collect substantive outcome data quickly. That’s the strongest differentiator the current SERP lacks — no ranking page shows outcome data. Testing scripts directly (sending messages) raises ethical/legal concerns and requires more operational complexity; it’s riskier and slower.
  - Higher trust & defensibility: A data-driven article with sample sizes, breakdowns, and lawyer-reviewed caveats is more defensible and link-worthy than another set of templates. It also leverages the staffer’s email threads as ground-truth validation.
  - Better long-term value: Readers will value quantified expectations; journalists and linkers (including paywalled NYT writers) may cite your data if it’s robust. Scripts without aggregated outcome evidence won’t beat Reddit’s anecdote advantage.

What evidence must exist before writing starts (concrete thresholds)
- Survey minimums and design
  - Response target: at least 400–600 complete responses from renters who have attempted to negotiate rent in the last 24 months. With 40k subscribers, a well-designed incentivized survey should meet that in 1–2 weeks.
  - Minimum subsample sizes: at least 50 responses each for the top tactics you’ll compare (e.g., asking at renewal, offering longer lease, using comps, offering to pay upfront). If some tactics have <50, report them as exploratory only.
  - Key variables to collect: location (city/state), landlord type (individual owner vs. management company), rent amount, change requested (absolute and percent), tactic(s) used, channel used (email/phone/in-person), timing relative to lease, whether they documented it in writing, outcome (reduced rent amount, partial concession, deferred payment, no change), time to outcome, and whether they can share redacted communication.
  - Outcome definition: a clear, single primary outcome (e.g., “any rent reduction agreed to in writing within 60 days”) and secondary outcomes (reply received, verbal concession, deferred payment).
- Validation and triangulation
  - At least 15–30 verifiable case submissions where respondents provide redacted emails/screenshots or allow us to view them; include the staffer’s 3 negotiation threads as verified exemplars.
  - 10–20 follow-up phone interviews to validate ambiguous responses and collect qualitative context.
- Statistical thresholds for claims
  - Be prepared to report confidence intervals and avoid overclaiming. For headline claims (e.g., “Tactic X is twice as likely to work as Y”), require p<0.05 and effect sizes that are practically meaningful (e.g., absolute difference ≥10 percentage points).
- Legal check
  - Lawyer retainer must review staffer threads and any template language; confirm no advice crosses into jurisdiction-specific legal counseling without disclaimers. Include a standard legal/disclaimer blurb reviewed by counsel.
- Editorial assets
  - Compelling visuals: charts showing success rates by tactic and landlord type; tables of scripts but only after verifying they correlate with better outcomes.
  - SEO and UX: clear methodology section (transparent survey instrument), and an FAQ addressing bias and limitations.

If you can meet those thresholds, write the Evidence-first page. If you cannot get 400–600 substantive responses or at least a dozen verifiable thread screenshots/interviews, pause: an evidence piece without sample size or verification risks being dismissed as anecdotal and won’t beat competitors.

Final recommendation
- Run a 10–14 day newsletter survey targeted to likely negotiators (filtering for renters who attempted negotiation in the past 24 months). Use a small incentive (raffle) and ask for optional redaction uploads and interview consent. If you hit the thresholds above, commission the article: data-driven hero section + methodology + actionable recommendations + lawyer-reviewed templates and links to redacted exemplar threads. If you can’t get the responses or verifications, don’t write another generic how-to — you’ll only add noise to a saturated SERP.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Find the honest gaps against a competitor page

Compare our page against the competitor page that outranks it. List what they cover that we do not, what they evidence that we merely assert (their tests, numbers, screenshots versus our claims), and what they do worse — thin sections, stale facts, padding — because the gap analysis must cut both ways to be trusted. Then name the three additions to our page that would most change what a reader gets from it, each tied to a specific gap you found. Do not recommend copying their structure; recommend closing evidence gaps. Work only from the two pages as provided. Our page: {{ours}} Theirs: {{theirs}}

Fill in: Both pages' content — full copy or honest section-by-section summaries.

Known limits: Cannot see authority signals (links, brand search) — a page can out-cover a competitor and still rank below it.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
OUR PAGE: /guides/cold-brew-at-home — sections: what cold brew is, ratio
(1:8, stated once), steep time (16-24h, no test data), grind size (coarse, one
sentence), a 4-step method, storage (up to 2 weeks, no source). ~900 words, no
images described, last touched 14 months ago.

THEIR PAGE (outranks us): sections: ratio table for concentrate vs ready-to-drink
(1:5 vs 1:8) with a photo of both side by side; steep-time experiment — same beans
at 12/16/20/24h with tasting notes per interval and a "16h wins for medium roasts"
verdict; grind comparison photos (coarse vs medium with extraction notes); method
for French press, mason jar, and dedicated maker; storage tested — flavor notes at
day 3/7/10 concluding "drink within a week, 2 weeks is technically safe but flat";
FAQ answering dilution and caffeine questions. Updated 2 months ago.

Model output (unedited)

Quick summary
- The competitor wins on evidence and sensory detail: they back recommendations with tests, photos, and measured timelines. Our page states correct basics but leaves nearly every practical claim unsupported.
- To be trusted (and to close the ranking gap) we should add targeted, evidence-focused content — not copy their layout, but replicate the kinds of evidence they supply.

1) What the competitor covers that our page does not
- Clear ratio table distinguishing concentrate vs ready-to-drink (1:5 vs 1:8) with a side‑by‑side photo.
- A steep‑time experiment (same beans at 12 / 16 / 20 / 24 h) with tasting notes and a verdict (“16 h wins for medium roasts”).
- Grind comparison with photos (coarse vs medium) and extraction notes.
- Multiple extraction workflows (French press, mason jar, dedicated maker) rather than a single 4‑step generic method.
- Storage testing with flavor notes at day 3 / 7 / 10 and a judged “drink within a week; 2 weeks technically safe but flat.”
- FAQ covering common practical questions (dilution, caffeine).
- Recent update (2 months ago).

2) What they evidence that we merely assert
- Ratio: They show a table and photo for concentrate vs ready‑to‑drink (1:5 vs 1:8). We state 1:8 once with no demonstration or clarification.
- Steep time: They ran a timed tasting experiment with notes and a specific verdict. We give a 16–24 h range with no test data or sensory evidence.
- Grind size: They provide comparison photos and extraction notes. We have a one‑sentence “coarse” recommendation with no visual or measured justification.
- Storage shelf life and flavor decline: They recorded flavor notes at multiple days and concluded a practical window. We claim “up to 2 weeks” with no testing or sensory timeline.
- Methods: They show how the brew changes by vessel/method; our single 4‑step method doesn’t address how method affects outcome or convenience.

3) What the competitor does worse (areas where their page is thin, possibly padded, or limited)
- Limited roast generalizability: Their steep‑time experiment gives a verdict for “medium roasts.” If they only tested one roast/profile, that result doesn’t necessarily apply to light or dark roasts — a limitation they don’t appear to caveat strongly.
- Potential method shallowness: Offering three methods (French press, mason jar, dedicated maker) is useful, but those sections can be shallow if they’re primarily photos + short steps without troubleshooting or timing differences. (From the summary you provided, depth per method isn’t stated.)
- Subjective language presented as practical guidance: “Drink within a week, 2 weeks is technically safe but flat” mixes sensory judgment and safety without explicitly separating flavor decline from food‑safety reasoning.
- Possible padding with photos: Photos (side‑by‑side ratio, grind shots) are helpful, but heavy reliance on images without quantified metrics (extraction yield, strength measurements, or scoring data) can look like visual polish over rigorous evidence.
- Narrow test scope implied: The steep‑time test appears to vary only time, not ratios, grind/time interaction, or multiple bean types — so readers might infer broader conclusions than the data support.

4) Three additions to our page that would most change what a reader gets (each tied to a specific gap)
Addition A — Short, reproducible steep‑time tasting test (closes the “we only assert steep time” gap)
- What to add: A compact experiment section: same coffee and grind, brew samples at 12 / 16 / 20 / 24 hours; include numeric ratings (e.g., 1–10) for acidity, sweetness, bitterness, body, and an overall score; include a one‑sentence verdict and a short caveat about roast dependency.
- Why it changes the page: It transforms our 16–24 h range from an unsupported assertion into an evidence‑backed recommendation readers can trust and replicate.
- Implementation notes (concise): Use exact coffee:water ratio (by weight), show sample preparation details (grind setting described in reference terms), and list the tasting results in a small, scannable block.

Addition B — Visual grind comparison + brief extraction notes (closes the “we assert coarse grind” gap)
- What to add: Two macro photos (coarse vs medium) at the same scale and a short explanation of how particle size affected extraction and taste in a quick test (e.g., coarse = cleaner/less bitter, medium = faster extraction/stronger). Give actionable guidance: recommended grind range for common methods (mason jar, French press).
- Why it changes the page: Readers making grind choices will immediately see and understand the effect rather than relying on a one‑sentence rule.
- Implementation notes: Keep it focused — two photos, one paragraph of sensory/functional comparison, and one line tying grind to method and steep time.

Addition C — Practical storage timeline with sensory notes and dilution guidance (closes the “we assert storage and ratio but give no evidence or clear dilution guidance” gap)
- What to add: A short tested storage timeline: flavor notes at day 3 / 7 / 10 (same protocol as steep test), a clear statement about safety vs freshness (e.g., flavor decline vs spoilage) and explicit dilution examples for concentrate vs ready‑to‑drink (show 1:5 concentrate and 1:8 ready‑to‑drink with quick examples: “If you brewed 1:5 concentrate, dilute 1 part concentrate to 4 parts water for drinking”).
- Why it changes the page: It converts vague “up to 2 weeks” advice into actionable expectations about how the brew will taste over time and how to prepare concentrate vs ready drinks.
- Implementation notes: If you can’t test microbiological safety, be explicit: separate sensory shelf‑life from safety guidance and recommend refrigeration and clean handling.

Final notes (about approach, not structure)
- Don’t copy the competitor’s exact page structure; instead, add compact, evidence‑focused blocks that slot into your current guide: a “Quick experiment” box for steep time, a “Grind at a glance” visual, and a “Storage & dilution” practical table. Each block should state methods/metrics clearly so readers know the scope and limits of the test.
- Because recency matters, update the page timestamp when these additions are published and note the scope of tests (beans used, roast level) so readers and search engines can judge applicability.

If you want, I can draft the exact wording and layout for each of the three additions (experiment protocol + sample tasting table, two grind photos with captions, and a storage/dilution mini‑table) tailored to your existing page tone.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Mine support questions for content topics

These are real questions customers asked support. Turn them into a content plan: group questions that are the same need in different words; for each group, judge whether it deserves a public page — recurring, answerable in general terms, and asked before purchase or early in use — or belongs in docs, or is account-specific noise; and for each page-worthy group, state the page's working title, the question phrasing to lead with (the customers' own words, which are usually the search phrasing too), and what the answer must include to actually close the question. Count how many tickets each group represents so the plan is ordered by evidence, not taste. Questions: {{questions}}

Fill in: Real questions from support tickets, chats, or sales calls — verbatim phrasing, one per line.

Known limits: Support questions skew to post-purchase problems — this surfaces retention content reliably, acquisition content only when pre-sales questions are in the sample.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-nano

Input fixture
QUESTIONS (from a meal-kit delivery service's support inbox, verbatim):
- can I skip a week without being charged
- how do I pause my deliveries while on vacation
- is the packaging recyclable? the ice packs specifically
- what do I do with the ice packs
- can you leave the box if I'm not home
- my box arrived warm is the food safe
- chicken arrived warm, can I still cook it
- do you have low sodium options
- can I swap an ingredient I'm allergic to
- how long do the meals keep in the fridge
- are your meals ok for gestational diabetes
- box was left in the sun all day
- when do I get charged each week
- can I change my delivery day
- why was I charged after I thought I cancelled

Model output (unedited)

Here’s a customer-request-driven content plan. I grouped the 15 questions into 7 recurring topics, judged whether each group should live on a public page (recurring, general, often asked before purchase or early use), or in docs (more internal or step-by-step) or as account-specific noise. For each public-page-worthy group I provide a working page title, the customer wording to lead with (the search-friendly phrasing you’ll likely see), and what the answer must include to actually close the question. I’ve ordered the plan by how many tickets each group has (evidence-driven).

Overall counts
- Total tickets: 15
- Groups with counts: 3, 3, 3, 2, 2, 1, 1
- Ordered by evidence (most tickets first): 3-ticket groups, then 2-ticket groups, then 1-ticket groups.

Group 1 (3 tickets) Billing cadence, skipping weeks, and cancellation policy
- Page-worthy? Yes (recurring, general, often asked pre- or early-use)
- Working page title: Billing cadence, skipping weeks, and cancellations
- Lead with customer wording (search-friendly phrases):
  - "can I skip a week without being charged"
  - "when do I get charged each week"
  - "why was I charged after I thought I cancelled"
- What the answer must include to close the question:
  - Clear explanation of how the weekly charge works (billing cycle, timing, and what events trigger charges)
  - How to skip a week and whether charges apply during a skip
  - How to change or cancel deliveries, including any cutoff times and how refunds or credits are handled if a cancellation occurs
  - What happens if a cancellation is processed late (refunds, partial charges, or next-cycle adjustments)
  - Step-by-step actions the user should take (in-app/website steps) and expected timelines
  - Contact or support flow if the user believes they were charged in error
  - A short FAQ subset for quick answers (e.g., “I skipped; why was I still charged?”)
- Why it belongs here: It addresses core billing and delivery timing questions that repeat across customers, both pre-purchase and during use.

Group 2 (3 tickets) Food safety when products arrive warm
- Page-worthy? Yes (recurring, safety-critical, generally asked before or during early use)
- Working page title: Box arrives warm or food-safe guidelines
- Lead with customer wording:
  - "my box arrived warm is the food safe"
  - "chicken arrived warm, can I still cook it"
  - "box was left in the sun all day"
- What the answer must include to close the question:
  - Immediate safety guidance: do not rely on smell or appearance; recommended immediate storage/handling steps
  - Temperature and time guidelines (e.g., if the product was exposed to heat beyond X hours, what is considered safe or unsafe)
  - Policy on replacements or refunds for temperature-affected shipments, including any required evidence (timestamps, photos)
  - Clear instructions for reporting safety concerns (where to contact, timeframe, required details)
  - Next steps for keeping food safe (thawing, refrigeration, reheating guidance if permissible)
  - Any preventive measures (shipping insulation, delivery timing) and what customers can do to minimize risk
- Why it belongs here: Temperature-related safety is a high-expectation, recurring concern; customers need a clear, public policy and rapid action steps.

Group 3 (3 tickets) Dietary options, substitutions, and suitability (including gestational diabetes and allergies)
- Page-worthy? Yes (recurring, general product features, often checked pre-purchase or early use)
- Working page title: Dietary options, substitutions, and allergy-aware features
- Lead with customer wording:
  - "do you have low sodium options"
  - "are your meals ok for gestational diabetes"
  - "can I swap an ingredient I'm allergic to"
- What the answer must include to close the question:
  - List of available dietary options (e.g., low-sodium, diabetic-friendly, other common diet accommodations)
  - Substitution policy: how to request ingredient substitutions, limits, timing/cutoffs, and any extra costs
  - Allergen handling: allergen cross-contact information, labeling practices, and how to note allergies in the order
  - Guidance on whether certain meals are guaranteed to be safe for gestational diabetes or allergies, including any labelling or certifications
  - How to modify future orders to reflect dietary needs, plus how and when changes take effect
  - Contact path for edge cases or special accommodations
- Why it belongs here: Nutrition/dietary needs are a common concern before purchasing and during ongoing use; customers frequently search for compatibility with health conditions/allergies.

Group 4 (2 tickets) Packaging, ice packs recycling and disposal
- Page-worthy? Yes (recurring, general, often checked pre-purchase or upon receipt)
- Working page title: Packaging and ice packs: recycling and disposal guidance
- Lead with customer wording:
  - "is the packaging recyclable? the ice packs specifically"
  - "what do I do with the ice packs"
- What the answer must include to close the question:
  - Recycling and disposal instructions for packaging materials and ice packs (what can be recycled, where to recycle, whether ice packs are reusable, and any required disposal steps)
  - Any composting guidance if applicable
  - Safety notes for handling cold packs (keeping away from children/pets, leakage precautions)
  - If there are options for returns/replacements or refunds related to packaging issues
  - Links to local recycling guidelines or provided packaging-specific disposal tips
- Why it belongs here: Packaging questions are frequent and practical; a public policy page reduces support load and helps customers dispose of materials correctly.

Group 5 (2 tickets) Delivery scheduling changes (delivery day and vacation pauses)
- Page-worthy? Yes (recurring, general, practical pre-use or early-use concern)
- Working page title: Change your delivery day and suspend deliveries
- Lead with customer wording:
  - "how do I pause my deliveries while on vacation"
  - "can I change my delivery day"
- What the answer must include to close the question:
  - How to pause deliveries (vacation hold) and the duration options
  - How to change the delivery day or frequency (cutoffs, blackout dates, impact on next shipments)
  - Any charges, refunds, or timing considerations when pausing or rescheduling
  - Steps to resume deliveries and verify the next shipment date
  - Troubleshooting tips and where to contact support if options aren’t visible
- Why it belongs here: Delivery timing and holds are common planning questions; a single public page helps customers manage schedules easily.

Group 6 (1 ticket) Delivery instructions: leave box if not home
- Page-worthy? Yes (recurring, general, often asked pre-purchase or first shipment)
- Working page title: Delivery instructions: leaving your box if you're not home
- Lead with customer wording:
  - "can you leave the box if I'm not home"
- What the answer must include to close the question:
  - Policy on leaving packages unattended: allowed locations (porch, with neighbor, secure drop point), any required authorization notes in the account
  - Any safety considerations (theft risk, weather risks)
  - How to update delivery instructions (where to add notes in the order, timing of changes)
  - What happens if no one is home (whether a second attempt is scheduled)
- Why it belongs here: Delivery instructions are a common, practical need; a public page reduces back-and-forth and aligns customer expectations.

Group 7 (1 ticket) Shelf life and storage
- Page-worthy? Yes (recurring, general, often checked pre-use)
- Working page title: Shelf life and storage of meals
- Lead with customer wording:
  - "how long do the meals keep in the fridge"
- What the answer must include to close the question:
  - Recommended fridge/storage durations for different meals
  - Storage best practices (temperature ranges, containers, freezing guidance if applicable)
  - Any product-specific caveats (e.g., opened ingredients, dressings)
  - Signs of spoilage and when to discard
  - Quick tips for maximizing freshness and safety
- Why it belongs here: Clear storage guidelines are a standard, high-value public resource.

Notes on implementation
- For each page-worthy group, consider a short FAQ layout plus a more detailed policy section. Use customer-friendly headings that align with common search queries.
- Include clear calls to action: “If you still need help after reading, contact us here” with a link to support.
- Cross-link relevant sections: e.g., a “Delivery and Billing” hub page could link to the individual delivery, billing, and pause/change pages.
- SEO considerations: use the exact customer phrases as headers or FAQ questions (as shown in the Lead with customer wording lines) to improve discoverability.
- Where to place content:
  - Public pages: Billing and Scheduling, Delivery Instructions, Temperature Safety, Dietary Options and Substitutions, Packaging & Ice Packs, Shelf Life.
  - Docs: If you have more granular procedures (e.g., internal escalation steps for safety incidents or refunds), keep those in docs or an internal knowledge base, but the high-level policy should be public.
  - Account-specific noise: None of these groups appear to require account-specific handling as the questions are general policy or product features; if a customer reports a unique billing issue, handle it via support ticket, but reference the public policy in the response.

Summary of recommended public pages (titles you can publish)
- Billing cadence, skipping weeks, and cancellations (lead phrases: “can I skip a week without being charged”, “when do I get charged each week”, “why was I charged after I thought I cancelled”)
- Box arrives warm or food safety: what to do (lead phrases: “my box arrived warm is the food safe”, “chicken arrived warm, can I still cook it”, “box was left in the sun all day”)
- Dietary options, substitutions, and allergen considerations (lead phrases: “do you have low sodium options”, “are your meals ok for gestational diabetes”, “can I swap an ingredient I’m allergic to”)
- Packaging and ice packs—recycling and disposal (lead phrases: “is the packaging recyclable? the ice packs specifically”, “what do I do with the ice packs”)
- Change delivery day and pause deliveries (lead phrases: “how do I pause my deliveries while on vacation”, “can I change my delivery day”)
- Delivery instructions: leave box if not home (lead phrase: “can you leave the box if I'm not home”)
- Shelf life and storage (lead phrase: “how long do the meals keep in the fridge”)

If you want, I can turn these into draft pages with fully fleshed-out copy blocks (intro, policy, steps, FAQs, and support links) ready to publish.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Outline a comparison page with fixed criteria

Outline an X-vs-Y comparison page where both options are judged on the same criteria throughout. First fix the criteria from what the provided audience evidence says actually drives the choice — not a generic feature list. Then outline: a verdict section that names who should pick which option and why; one section per criterion with both options addressed side by side; and a table plan listing rows and what fills each cell, marking any cell needing hands-on verification versus vendor-published fact. Forbid a structure that reviews X fully then Y fully — readers comparing need the same-criterion adjacency. Options and audience evidence: {{context}}

Fill in: The two options, who is choosing between them, and what you know about what drives the decision.

Known limits: Criteria fixed from thin audience evidence drift generic — real forum questions and support tickets beat guessed criteria.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-nano

Input fixture
OPTIONS: Shopify vs WooCommerce, for a web-dev agency's blog whose
readers are small-business owners with no developer on staff.

AUDIENCE EVIDENCE (from the agency's sales calls and forum research):
- Most-asked question by far: "what will this actually cost me per month once
  I'm running" — including apps/plugins, not the base price
- Second: "can I leave later without losing everything"
- Frequent pain: WooCommerce owners arriving after a hack or a broken update
  with no one to call
- Rare: nobody asks about theme quantity or design flexibility unprompted
- Payment processing fees confuse everyone; Shopify's extra fee for external
  gateways is routinely discovered late

Model output (unedited)

Here’s an outline for an X-vs-Y comparison page (Shopify vs WooCommerce) that uses the same criteria for both options, driven by the audience evidence you provided. It centers on what actually matters to small-business owners with no developer on staff, not a generic feature list.

1) Criteria fix: what actually drives the choice (based on audience evidence)
- Total monthly cost including apps/plugins (not just base price)
  - What the audience asks most: “What will this cost per month once I’m running, including apps/plugins?”
  - What to measure: base platform price, add-on/app costs, gateway/processing fees, hosting (for WooCommerce), and any other recurring costs.
- Ability to leave later without losing everything (data portability)
  - Audience worry: “Can I leave later without losing everything?” including product data, customers, orders, media.
  - What to measure: data portability, exports/imports, migration support, data ownership.
- Security risk and support when something goes wrong
  - Frequent pain: WooCommerce owners facing hacks or broken updates with no one to call.
  - What to measure: built-in security/patch cadence, availability of support, incident response, backups.
- Hosting/maintenance burden and responsibility
  - For WooCommerce: self-hosted, you manage hosting, backups, security, updates.
  - For Shopify: hosted by the platform; you rely on their patching and infrastructure.
- Ease of use for non-developers
  - Audience profile: small-business owners with no staff developer.
  - What to measure: onboarding, day-to-day management, app ecosystem, learning curve.
- Payment processing costs and gateway behavior
  - Rare but real: Shopify’s extra fee for external gateways, often discovered late.
  - What to measure: whether gateway fees apply, and how much they actually add to monthly costs.
- (Note on design flexibility) Design/theme flexibility is relatively rare as a driver per your evidence; can be deprioritized but mention briefly as a secondary factor.

2) Verdict (who should pick which option and why)
- Overall guidance for small-business owners with no developer on staff:
  - Shopify is typically the better fit for most readers in this group due to predictable monthly costs, managed hosting, built-in security, simpler setup, and strong support. It reduces the risk of hacks and broken updates and minimizes the need for a technical person to keep things running.
  - WooCommerce is worth it if you already have WordPress hosting, want to avoid recurring platform fees, have some technical ability, and can tolerate more hands-on maintenance. It can offer lower ongoing costs only if you can manage hosting, security, updates, and backups yourself and you’re willing to handle more complexity.
- Quick one-liner:
  - If predictability, security with minimal effort, and ease of use for non-technical staff matter most, go Shopify.
  - If you already have WordPress hosting and a technical resource (or you’re comfortable managing hosting and security) and want maximum control and potentially lower recurring platform costs, go WooCommerce.

3) One section per criterion (same-criterion adjacency; side-by-side for Shopify and WooCommerce)
- Criterion: Total monthly cost including apps/plugins
  - Shopify: base plan plus apps; possible gateway fees for non-Shopify Payments; costs tend to be predictable but can rise with apps.
  - WooCommerce: no core subscription; hosting costs vary; plugin/add-on fees vary; gateway fees depend on chosen processors; total can be lower or higher depending on setup.
  - Quick note: numbers depend on shop size and apps; actual totals require live pricing.

- Criterion: Data portability / exit
  - Shopify: exports available (customers, orders, products) but moving off can involve friction and might require apps; some data dependencies (media, configurations) can complicate full migration.
  - WooCommerce: data lives in WordPress/database; exports/imports are straightforward; generally easier to migrate to another platform or back up locally.
  - Quick note: test export/import with your own data to confirm ease of migration.

- Criterion: Security and support
  - Shopify: managed hosting with built-in security, PCI compliance, automatic updates; 24/7 support; less reliance on an internal tech person.
  - WooCommerce: security/patching depends on your hosting provider and plugin maintenance; you’ll need to manage backups, updates, and potential security hardening; support is typically per-plugin or hosting.
  - Quick note: real-world risk depends on hosting quality and how you manage updates.

- Criterion: Hosting/maintenance burden
  - Shopify: fully hosted; Shopify handles hosting, backups (per plan), uptime, and performance.
  - WooCommerce: self-hosted; you choose hosting, run backups, manage updates, security, and performance; more control but more work.
  - Quick note: “maintenance” is a constant in WooCommerce; “hands-off” in Shopify.

- Criterion: Ease of use for non-developers
  - Shopify: designed for non-technical users; streamlined admin, app ecosystem, guided setup.
  - WooCommerce: relies on WordPress admin; can be less intuitive; more moving parts (themes, plugins) to manage.
  - Quick note: onboarding time is typically shorter on Shopify.

- Criterion: Payment processing costs and gateway behavior
  - Shopify: Shopify Payments usually eliminates additional gateway fees; using third-party gateways can trigger extra transaction fees on some plans; many readers discover this late.
  - WooCommerce: no platform-level gateway fees; payment costs depend on chosen processors (Stripe, PayPal, etc.) and any gateway plugin costs.
  - Quick note: confirm current policy for your plan and gateways.

- Criterion: Design flexibility (secondary driver)
  - Shopify: solid design options with themes; some limitations on deep code changes without apps.
  - WooCommerce: broad design control via WordPress themes and page builders; higher flexibility but more complexity.
  - Quick note: largely a secondary factor per evidence; not the main driver for most readers.

4) Table plan (rows, cells, and verification notes)
Rows (criteria), Columns (Shopify, WooCommerce), plus notes on content type and verification needs
- Row 1: Total monthly cost including apps/plugins
  - Cell (Shopify): Summary of base plan + apps; potential external-gateway fees; vendor-published pricing guidance; numbers require hands-on verification for your shop.
  - Cell (WooCommerce): Summary of hosting cost + plugin/add-on fees + gateway costs; vendor-published guidance on core “WooCommerce is free” but total depends on setup; numbers require hands-on verification.
- Row 2: Data portability / exit
  - Cell (Shopify): Data export capabilities; some migration friction; vendor-published facts; hands-on verification recommended to test your data and assets.
  - Cell (WooCommerce): Data portability via WordPress export/import; generally straightforward; vendor-published facts; hands-on verification optional but helpful.
- Row 3: Security and support
  - Cell (Shopify): Managed security, 24/7 support; vendor-published facts; hands-on verification optional but can confirm response times for your region.
  - Cell (WooCommerce): Security depends on hosting and plugins; vendor-published facts; hands-on verification recommended to assess your actual hosting security and support coverage.
- Row 4: Hosting/maintenance burden
  - Cell (Shopify): Hosted solution; minimal maintenance; vendor-published facts; hands-on verification optional but can confirm perceived effort.
  - Cell (WooCommerce): Self-hosted; you manage hosting, backups, updates; vendor-published facts; hands-on verification recommended for your host and backup strategy.
- Row 5: Ease of use for non-developers
  - Cell (Shopify): User-friendly for non-devs; vendor-published facts; hands-on verification optional to confirm real-world onboarding with your team.
  - Cell (WooCommerce): WordPress admin—more complexity; vendor-published facts; hands-on verification recommended to test a sample setup.
- Row 6: Payment processing costs and gateway behavior
  - Cell (Shopify): External gateway fees if not using Shopify Payments; vendor-published facts; hands-on verification advisable to calculate for your expected gateways.
  - Cell (WooCommerce): No platform-wide gateway fees; costs depend on gateway and plugins; vendor-published facts; hands-on verification recommended for your chosen processors.
- Row 7: Design flexibility (secondary)
  - Cell (Shopify): Themes and limited code access; vendor-published facts; hands-on verification optional.
  - Cell (WooCommerce): Greater design control via WordPress themes; vendor-published facts; hands-on verification optional.

Notes on how to populate the table
- For each cell, aim to present a concise, side-by-side summary of how Shopify and WooCommerce compare on that criterion.
- Mark any cell with “Hands-on verification required” when real-world numbers or behavior depend on your exact shop configuration (e.g., exact monthly cost with chosen apps, specific hosting plan, exact gateway fees, migration steps with your data).
- Mark cells as “Vendor-published fact” when you can cite official pricing or policy statements (e.g., “Shopify hosting is included in plan,” “WooCommerce core is free,” “Shopify Payments reduces gateway fees on supported plans,” etc.).
- Provide a short “What to test” note for hands-on verifications (e.g., simulate a 1-month plan with typical apps, run a data export, run a security backup, request a migration demo).

If you want, I can turn this into a fillable page skeleton with placeholders for yours numbers and a suggested layout that enforces the same-criterion adjacency (verdict at top, then Criterion sections, then a universal table).

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Outline a page from the SERP's dominant intent

From this SERP evidence, first state the dominant intent in one sentence and name which results prove it. Then outline a page that serves that intent: headings in searcher order — the committed answer early, supporting detail after, edge cases last — with one line per section stating what it must establish and what evidence it needs. Where the SERP shows a secondary intent worth serving on the same page, mark that section as serving it; where the secondary intent needs its own page, say so instead of bloating this one. Keyword: {{keyword}} SERP: {{serp}}

Fill in: The keyword and an annotated SERP: what ranks and what each result is.

Known limits: An outline shaped by today's SERP inherits today's SERP — revisit if the results change class (e.g. products replace guides).

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
KEYWORD: raised garden bed soil mix

SERP (annotated):
1. "The Perfect Soil Mix for Raised Beds" - GrowVeg — recipe: 1/3 compost, 1/3 topsoil,
   1/3 aeration; explains each component's job
2. "Raised Bed Soil Calculator" - GardenMath — interactive cubic-feet calculator, thin
   text around it
3. "What I Put in My Raised Beds (Cheap)" - YouTube, 12 min — budget fill methods,
   hugelkultur base layer
4. "Mel's Mix vs Native Soil: 2-Year Comparison" - PlotNotes — actual side-by-side
   yield data
5. "Raised Bed Soil" - Lowe's category page — bagged products
People Also Ask: "How many bags of soil for a 4x8 raised bed?", "Can I use only
compost?", "What is the cheapest way to fill a raised bed?"

Model output (unedited)

Dominant intent (one sentence): Users want a clear, practical recipe and buying/how‑much guidance for raised bed soil — proven by the recipe article (GrowVeg #1), the how‑many calculator (#2), budget fill methods video (#3), performance comparison (#4), and product listings (#5).

Page outline (headings in searcher order — committed answer first, supporting detail after, edge cases last). Each line: heading — what it must establish — evidence it needs (which SERP results prove or supply that evidence). Marked sections that also serve secondary intents.

1) H1 / Quick answer: "Best raised bed soil mix — the simple recipe and how many bags you need"
   - Must establish the committed recommendation (exact mix ratio), a one‑line why, and a one‑click link to the calculator and buy options.
   - Evidence needed: recipe authority and rationale (GrowVeg #1, PlotNotes #4); calculator link (GardenMath #2); product links (Lowe's #5).

2) "The recipe: 1/3 compost, 1/3 topsoil (or loam), 1/3 aeration (perlite/coarse sand) — and alternatives"
   - Must establish the standard Mel’s Mix / balanced formula, define each component, and list acceptable substitutes (e.g., peat vs coir, perlite vs pumice).
   - Evidence needed: recipe and component roles (GrowVeg #1); alternative mixes and outcomes (PlotNotes #4).

3) "How much do I need? Instant calculator + quick bag estimates" (serves secondary intent: tools/purchase planning)
   - Must provide an embedded/linked cubic‑feet calculator, example answers for common beds (e.g., 4×8×1ft), and conversion to number of common bag sizes and cost ballpark.
   - Evidence needed: interactive calculator (GardenMath #2); common bag volumes/prices (Lowe's #5); PAA question "How many bags for a 4×8?" for examples.

4) "Buy vs mix: costs, pros/cons, and what to buy at the store" (serves secondary intent: transactional/shopping)
   - Must establish when to buy premixed bagged soil vs mix your own, give a shopping list (how many bags, bulk topsoil, compost sources), and include links to products.
   - Evidence needed: product availability and bag sizes/prices (Lowe's #5); cost‑saving hints and context (YouTube #3).

5) "Cheapest ways to fill a deep bed (shortcuts and budget builds)" (serves secondary intent: cheapest-fill advice)
   - Must establish safe low‑cost methods (hugelkultur base, partial native soil plus compost, layered cheap fills), risks and expected longevity/yield.
   - Evidence needed: budget fill methods (YouTube #3); practical outcome comparisons (PlotNotes #4).

6) "Performance: Mel’s Mix vs native soil (what the data shows)"
   - Must establish yield/health differences, timelines (how many seasons to see effects), and when native‑soil amendment is acceptable.
   - Evidence needed: side‑by‑side yield data and analysis (PlotNotes #4); supporting rationale from recipe article (GrowVeg #1).

7) "Can I use only compost? Quick answers to common questions" (serves secondary intent: PAA)
   - Must answer PAA items concisely (only compost? how many bags for 4×8? what’s the cheapest way?), and point to calculator, budget section, and caveats.
   - Evidence needed: compost role and limits (GrowVeg #1); calculator examples (GardenMath #2); budget options (YouTube #3).

8) "Step‑by‑step: mixing, filling, and first season care" (supporting detail)
   - Must establish a concise how‑to: mixing ratios by volume, layering technique for deep beds, watering/crop selection for year one.
   - Evidence needed: recipe and practical tips (GrowVeg #1, YouTube #3).

9) "Troubleshooting & edge cases" (last)
   - Must cover dense clay sites, limited compost availability, shallow beds, reusing old raised‑bed soil, and soil testing recommendations — provide decision criteria for altering the mix.
   - Evidence needed: guidance from recipe and comparison pieces (GrowVeg #1; PlotNotes #4); practical fixes from budget methods (YouTube #3). May need additional authoritative citations beyond given SERP for technical fixes.

10) "Further reading / tools (separate pages recommended)"
    - Must identify secondary intents that warrant their own page instead of bloating this one: a full ecommerce product‑comparison and price‑tracker for bagged mixes (needs Lowe’s‑style catalog + live prices), and a deep DIY page on hugelkultur and deep‑fill techniques (needs step‑by‑step plans and long‑form video/tutorials).
    - Evidence needed: product listings (Lowe's #5) show transactional depth; YouTube #3 indicates long‑form how‑to content better as its own dedicated page.

Use this structure to deliver the fast answer up top (1), an actionable calculator/shop path (3–4), tested alternatives and data (5–6), then how‑to details (8) and edge cases (9); include short PAA Q&A (7) for searchers looking for quick answers.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Plan a quarter's content around known dates

Plan one quarter of content for this site around the fixed dates provided — seasonal peaks, industry events, recurring deadlines. For each planned piece: the event it serves, the publish date that beats the demand curve (pages need indexing and settling time before the searches arrive, so state the lead time you are assuming), whether it is a new page or a refresh of the listed existing page, and what must be true for it to be ready. Use only the events and pages provided — do not invent search volumes or events. Mark the pieces that are evergreen-with-a-seasonal-spike versus purely seasonal, because the former deserve better URLs (no year in the slug). Site, events, existing pages: {{context}}

Fill in: What the site covers, the quarter's known dates and events, and existing seasonal pages you already have.

Known limits: Plans around the events you listed — a competitor planning around events you forgot still beats you to them.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-nano

Input fixture
SITE: content site for home vegetable gardeners, US audience,
strongest in zones 5-7

QUARTER: January-March

KNOWN DATES/EVENTS: seed catalogs arrive Dec-Jan; indoor seed-starting for
zones 5-7 runs Feb-Mar (8-10 weeks before last frost, typically mid-Apr to
mid-May); many extension offices publish updated frost maps in Jan; St.
Patrick's Day = traditional potato-planting marker; spring equinox Mar 20.

EXISTING PAGES: /guides/seed-starting-indoors (solid, updated last year),
/guides/when-to-start-tomatoes (thin, 2023), /blog/best-seed-catalogs-2024
(year in slug, decayed), /guides/last-frost-date-lookup (healthy traffic
year-round, spikes Feb-Apr).

Model output (unedited)

Here’s a focused content plan for January–March (quarter) aligned to the fixed dates/events you provided. I’ve tied each planned piece to a specific event, set a publish date that beats the demand curve (with assumed indexing/settling lead time), and noted whether it’s a new page or a refresh of an existing page. I also flag evergreen-with-seasonal-spike vs purely seasonal and list what must be ready to publish.

Planned pieces

1) Best seed catalogs (refresh; evergreen-with-seasonal-spike)
- Event served: Seed catalogs arrive Dec–Jan (seasonal spike; evergreen relevance thereafter)
- Proposed slug: /guides/best-seed-catalogs
- Publish date (lead time): Jan 15 (lead time ~14 days before peak late Jan/early Feb interest)
- Page type: Refresh of /blog/best-seed-catalogs-2024
- Readiness criteria:
  - Update list with current-year seed catalogs and verify links
  - Remove year from slug and implement 301 redirect from old URL
  - Add brief intro about how catalogs influence early planning for zone 5–7
  - Update visuals and any affiliate links; ensure SEO metadata
- Why this for Jan: aligns with catalog season and serves readers planning early-season purchases
- Evergreen vs seasonal: Evergreen-with-seasonal-spike

2) Indoor seed-starting Indoors (refresh; evergreen-with-seasonal-spike)
- Event served: Indoor seed-starting for zones 5–7 runs Feb–Mar (8–10 weeks before last frost)
- Proposed slug: /guides/seed-starting-indoors
- Publish date (lead time): Jan 28 (lead time ~12–14 days)
- Page type: Refresh of /guides/seed-starting-indoors
- Readiness criteria:
  - Update timing and steps for zones 5–7 (Feb–Mar starting windows)
  - Integrate any updated frost-date guidance (link to Last Frost Date Lookup)
  - Add current best practices for starting tomatoes and peppers indoors (if relevant)
  - Fresh images or diagrams showing seed-start setup
  - SEO check: optimize for “indoors seed starting zones 5-7,” etc.
- Why this for Jan: supports the core indoor-starting window and capitalizes on the February surge
- Evergreen vs seasonal: Evergreen-with-seasonal-spike

3) When to start tomatoes (refresh; purely seasonal with evergreen base)
- Event served: Indoor start timing for tomatoes; supports spring planting schedule
- Proposed slug: /guides/when-to-start-tomatoes
- Publish date (lead time): Feb 8 (lead time ~10–14 days)
- Page type: Refresh of /guides/when-to-start-tomatoes (2023)
- Readiness criteria:
  - Update tomato-start timelines for zone 5–7 using current frost map guidance
  - Expand to include brief indoor-start schedule, transplant window, and outdoor hardening tips
  - Improve structure and add more actionable steps (checklists by week)
  - Ensure internal links to seed-starting indoors and frost-date lookup
- Why this for Feb: fresh guidance aligns with ongoing indoor-starting and pre-frost planning
- Evergreen vs seasonal: Purely seasonal (timing shifts each year; not evergreen slug)

4) Last frost date lookup (refresh; evergreen-with-seasonal-spike)
- Event served: Frost maps updated in Jan; spikes Feb–Apr
- Proposed slug: /guides/last-frost-date-lookup (or keep /guides/last-frost-date-lookup if you prefer)
- Publish date (lead time): Feb 1 (lead time ~14 days)
- Page type: Refresh of /guides/last-frost-date-lookup
- Readiness criteria:
  - Incorporate January frost-map updates for zones 5–7
  - Verify the interactive map/tool works across devices
  - Add quick-start planting windows by crop (especially for indoor-to-outdoor transitions)
  - Update related content links (seed-starting indoors, tomato timing)
- Why this for Feb: preps gardeners for the Feb–Apr spike in frost-date searches
- Evergreen vs seasonal: Evergreen-with-seasonal-spike

5) St. Patrick’s Day potato planting guide (new page; purely seasonal)
- Event served: St. Patrick’s Day as a traditional potato-planting marker (mid-March)
- Proposed slug: /guides/potato-planting-st-patricks-day
- Publish date (lead time): Feb 28 (lead time ~14 days)
- Page type: New page (not a refresh of an existing page)
- Readiness criteria:
  - Provide region-appropriate potato plantings for zones 5–7
  - Include recommended varieties suitable for early planting and common pests/diseases
  - Step-by-step planting guidelines, conditioning soil, and basic care
  - Cross-link to frost-date lookup and seed-starting resources
  - Visuals or diagrams showing potato planting steps
- Why this for Feb: aligns with a traditional marker and captures March search interest
- Evergreen vs seasonal: Purely seasonal

6) Spring equinox planting calendar (new page; purely seasonal)
- Event served: Spring equinox (Mar 20) as a planting milestone
- Proposed slug: /guides/spring-equinox-planting-calendar
- Publish date (lead time): Mar 6 (lead time ~14 days)
- Page type: New page
- Readiness criteria:
  - Create a clear zone-5–zone-7 planting calendar around the equinox
  - Include recommended crops and sowing windows that align with last-frost estimates
  - Cross-link to seed-starting indoors and last frost date lookup
  - Include a simple downloadable calendar or printable version if possible
- Why this for Mar: captures readers planning for the start of the season around the equinox
- Evergreen vs seasonal: Purely seasonal

Notes on execution and readiness
- Slug strategy: For evergreen-with-seasonal-spike items, use year-less slugs (no year in URL) to support long-term ranking. If you must keep older posts, implement 301 redirects from dated URLs.
- Interlinking: Each plan should link to /guides/last-frost-date-lookup and /guides/seed-starting-indoors to reinforce the seasonal planning workflow.
- Asset needs: Ensure up-to-date visuals, checklists, and any charts are recreated or refreshed; verify all links (seed catalogs, varieties, maps) are current.
- SEO and metadata: Add or refresh meta titles and descriptions to reflect current year (where appropriate for time-sensitive pages) while maintaining evergreen value.
- Analytics: Track performance for refreshed pages vs new pages to adjust promotion in subsequent weeks.

Summary timeline (planning view)
- Jan 15: Publish/refresh Best seed catalogs (evergreen-with-seasonal-spike)
- Jan 28: Publish/refresh Indoor seed-starting (evergreen-with-seasonal-spike)
- Feb 1: Publish/refresh Last frost date lookup (evergreen-with-seasonal-spike)
- Feb 8: Publish/refresh When to start tomatoes (purely seasonal)
- Feb 28: Publish St. Patrick’s Day potato planting guide (purely seasonal)
- Mar 6: Publish Spring equinox planting calendar (purely seasonal)

If you’d like, I can draft suggested copy blocks (intro, section headers, and a short CTA) for each piece and propose concrete slug adjustments (including suggested canonical redirects) to ensure a smooth rollout.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Plan a statistics page writers will cite

Plan a statistics page for this niche that earns citations because its numbers are original. Identify: which questions in this niche get asked with no current, sourced number attached; which of those this operation could genuinely answer from its own data or a survey it could realistically run; and for each chosen statistic, the exact methodology note that must accompany it — sample, period, collection method — because an uncited method is why stats pages fail. Plan the update cadence and what gets versioned when numbers change. Reject statistic ideas that would repackage others' numbers; aggregation is not the asset, originality is. Niche and available data: {{context}}

Fill in: Your niche, what proprietary data you sit on, and what you could realistically collect.

Known limits: The plan assumes your data access is as described — overstated access produces a plan you cannot execute.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
NICHE: wedding photography business advice (site read by working
wedding photographers)

AVAILABLE DATA:
- The site runs a job board where photographers post second-shooter gigs with
  day rates visible: ~3,100 postings over 4 years, tagged by US region
- Annual reader survey (1,200-1,800 responses each year, 5 years running) asks:
  packages offered, average booking price, weddings shot per year, editing
  hours per wedding, insurance carried
- Could add questions to next survey (fields in March)
- No access to: platform booking data, industry-wide revenue figures

Model output (unedited)

Goal: build a statistics page for working wedding photographers that earns inbound citations because every published number is original, reproducible, and accompanied by an explicit methodology note. Below I (1) list common unsourced questions in the niche, (2) pick the subset your operation can realistically answer with its own data or a realistic survey add-on, (3) give an exact methodology note for each published statistic (sample, period, collection method, inclusion/exclusion/cleaning rules), (4) define update cadence and versioning rules, and (5) reject statistics you should not attempt because they would merely repackage third‑party numbers or require unavailable platform data.

1) Common questions people ask with no reliable, sourced number
- “What does a second shooter get paid, regionally?”
- “What’s the typical booking price for a full wedding day?”
- “Do most pro wedding photographers sell packages or a la carte?”
- “How many weddings does a working photographer shoot per year?”
- “How many editing hours does a wedding take on average?”
- “How many photographers carry business insurance?”
- “What deposit percentage is normal?”
- “How long until clients receive the final gallery?”
- “How often do main shooters hire second shooters and what do they pay?”
- “How many photographers outsource editing and how much do they pay per wedding?”
- “How often do photographers raise prices and by how much?”
- “What’s the typical contract cancellation/refund policy?”

Many of these circulate as “conventional wisdom” but lack a current, citable measurement tied to a clear sample and method.

2) Which of those you can answer credibly (and how)
You can credibly answer the following using data you already have or by adding a small number of targeted questions to your next annual survey (March). I mark which data source to use.

A. Second-shooter day rates by US region — job board data (existing)
B. Distribution of “average booking price” (main shooter) — annual reader survey (existing)
C. Percent offering packages vs à la carte — annual reader survey (existing)
D. Weddings shot per year (distribution) — annual reader survey (existing)
E. Editing hours per wedding (median & distribution) — annual reader survey (existing)
F. Percent carrying business insurance — annual reader survey (existing)
G. Typical deposit percentage — add to next survey (realistic)
H. Typical final-gallery delivery time (days) — add to next survey
I. Frequency main shooters hire second shooters (per year) — add to next survey
J. Percent outsourcing editing + average price paid per outsourced wedding — add to next survey
K. Price-increase cadence and typical percent increase — add to next survey

Notes on feasibility:
- Job-board postings are already numeric and region-tagged; that makes A immediately publishable.
- You already collect B–F annually with large n (1,200–1,800) so these are straightforward.
- G–K require adding 6–8 short, tightly worded questions to the March survey; response rate should be comparable to other items.

3) Exact methodology note for each chosen statistic
Below are the exact methodology notes to publish with each statistic. Use the language verbatim on the site so journalists and researchers can cite it.

A. Median second‑shooter day rate by US region
- Sample: All job-board postings that include a numeric day rate and a US region tag, posted between 2022‑01‑01 and 2025‑12‑31 (inclusive). Current dataset size: n = 3,100 postings.
- Collection method: Automated scrape of job-board posting form fields; region is the poster-selected region tag. Numeric extraction rules: if posting gives a range (e.g., $150–$250), use the mid-point; if it lists an hourly rate only, exclude. Currency: U.S. dollars as posted.
- Cleaning & exclusions: Remove duplicates by (identical poster ID + identical text + posted within 7 days); remove postings with day rates < $50 or > $2,000 as likely input error; keep unpaid/trade entries only if a numeric $ value appears.
- Metrics reported: median, 25th & 75th percentiles, sample size per region, and number of excluded postings. Do not annualize or attempt to convert to “per wedding” rates.
- Caveats: This is a marketplace posting dataset (what posters offered to pay), not a record of completed hires.
- Minimum cell reporting: only report regional breakdowns where n ≥ 30 postings.
- Update cadence: update this table monthly (append new postings) and publish quarterly snapshots. Every publish includes the period covered and n.

B. Distribution of average booking price (main shooter)
- Sample: Respondents to the site’s annual reader survey who (a) identify as a primary/lead wedding photographer and (b) answered the “what is your average booking price?” question. Use the most recent survey year for the current snapshot; show a 5‑year trend using prior annual surveys (2019–2023).
- Period: Respondents answer based on their average booking price for contracts signed in the prior 12 months.
- Collection method: Self‑reported numeric entry (round to nearest dollar). For historical trend, use the same numeric field from earlier years.
- Cleaning & exclusions: Exclude entries < $100 or > $100,000 as likely miscoded; where respondents left blank, exclude. Report both median and mean, plus 10th/90th percentiles.
- Metrics reported: median, mean, IQR, n for that year, and 5‑year trend line with n per year.
- Weighting: none (raw respondents). Report survey response count and margin-of-error for proportions where applicable.
- Caveats: Self‑reported and likely skewed toward readers of this site (working pros, English‑speaking). Sample frame: your readership.
- Update cadence: Annual (after each survey). Archive each year’s dataset and keep historic snapshots.

C. Percent offering packages vs a la carte
- Sample/period: Same as B; use most recent annual survey; include only respondents who identify as active lead photographers.
- Collection method: Survey question with multiple-choice (Packages only; Mostly packages with a la carte options; Even mix; Mostly a la carte; A la carte only).
- Metric: percentage in each category, n, 95% binomial confidence interval for each proportion.
- Cleaning: Exclude respondents who indicate they are not actively booking weddings.
- Update cadence: Annual.

D. Weddings shot per year (distribution)
- Sample/period: Annual survey respondents who identify as active wedding photographers and answered “how many weddings did you shoot in the past 12 months?” Use most recent year for snapshot; supply 5‑year trend.
- Collection method: Numeric entry (integer). Include options for part‑time shading (e.g., 0–5, 6–15, etc.) if needed.
- Cleaning: Exclude implausible entries (> 500 weddings/year). Report median, mean, percentiles, and share categories (0–10, 11–25, 26–50, 51+).
- Update cadence: Annual.

E. Editing hours per wedding
- Sample/period: Annual survey respondents reporting average editing hours per wedding in the past 12 months.
- Collection method: Numeric entry (hours). Provide definition in question: “include color grading, culling, basic retouching of delivered images; exclude deliverable creation (albums) and client communications.”
- Cleaning: Exclude values > 200 or < 0.5 as likely errors. Report median, mean, IQR, and the share outsourcing any editing.
- Update cadence: Annual. Consider splitting by package level (e.g., full-day vs elopement) if n allows.

F. Percent carrying business insurance
- Sample/period: Annual survey item asking “Do you carry business insurance (liability, equipment, both, none)?”
- Collection method: Multiple choice. Report percent in each category with n and 95% CI.
- Update cadence: Annual.

G. Typical deposit percentage (new survey question)
- Question wording to add in March: “What percentage of the total contract fee do you typically require as a deposit to secure a booking? Enter a number between 0 and 100.”
- Sample: All respondents who identify as lead wedding photographers and who booked at least one wedding in the past 12 months.
- Period: responses collected March (current year).
- Cleaning/exclusions: Exclude entries <1 or >100. If multiple deposit practices stated in text, pick the most common (survey should force single numeric answer).
- Metrics: median, mean, IQR, percent using non-refundable deposits (ask follow‑up: “Is your deposit refundable if the client cancels? yes/no”).
- Minimum reporting n: 200 respondents; otherwise flag low-n.

H. Typical final-gallery delivery time (days) (new survey)
- Question wording: “After a wedding date, how many calendar days do clients typically wait to receive the final edited gallery (full delivery)? Enter an integer number of days.”
- Sample/period/cleaning: same eligibility as G. Exclude entries >365 unless they explain seasonal backlog; report median, percent ≤30 days, ≤60 days, etc.
- Update cadence: Annual.

I. How often main shooters hire second shooters (new survey)
- Question wording: “In the past 12 months, on how many weddings did you hire a second shooter?” (numeric) + “When you hire, do you pay per day or per wedding? (day/wedding/both)”
- Sample & cleaning: include only active lead photographers. Report distribution (median hires/year), percent who never hire, and percent who pay per-day vs per-wedding.
- Update cadence: Annual.

J. Percent outsourcing editing and average price paid per outsourced wedding (new survey)
- Question wording (two items): 1) “Do you outsource any portion of your wedding editing? (No / Some / All).” 2) If Some/All: “What is your average cost paid to a third-party editor per wedding? $____”
- Sample: active photographers who edit any portion or outsource.
- Cleaning: exclude outliers >$5,000; report median and IQR, plus share outsourcing none/some/all.
- Update cadence: Annual.

K. Price-increase cadence and percent raise (new survey)
- Question wording (two items): 1) “How often do you increase your prices? (Never / Every 1 year / Every 2 years / Every 3+ years / Other (specify)).” 2) “When you increase prices, what is the typical percent increase? Enter a number 0–200.”  
- Sample & cleaning: active lead photographers. Report distribution and median percent increase.
- Update cadence: Annual.

4) Additional data & survey design recommendations (exact question wording you can copy)
- Gate reliably: start survey with “Are you an active lead wedding photographer who has booked at least one wedding in the past 12 months?” Only allow eligible respondents to answer business questions.
- Force numeric fields where you need numeric analysis (no text ranges). Provide ranges only as backup for those who prefer to select ranges.
- Keep new items short — add 6–10 targeted questions maximum in March to avoid response fatigue.
- For sensitive items (revenue), prefer binned categories (easier to answer and higher response rate). If you want revenue later, ask as bins, not exact numbers.

5) Update cadence, versioning, and publication policy
- Two update cadences:
  - Job‑board metrics (A): continuous ingestion, monthly refreshes; publish a time‑stamped quarterly snapshot page (e.g., “Job Board Snapshot — Q2 2026”). Each snapshot must list period covered and n.
  - Survey metrics (B–F plus G–K): annual refresh tied to the annual survey. Publish as “Annual Photographer Business Survey — 2026” with data and a historical comparison chart (2019–2026).
- Versioning rules:
  - Each published metric gets a version token: [metric] vYYYY.MM.DD (date of publication). Example: “Second-shooter median rate — v2026.08.01”.
  - Maintain an audit/change log page listing: metric name, old value, new value, publish date, n, and whether methodology changed. If only data changed but method stayed the same, label it “data update”. If methodology changes (e.g., changed exclusion thresholds), increment major version and document exactly what changed and why.
  - Keep and publish raw anonymized aggregate CSVs for each public snapshot (columns: metric, value, n, period, note), and store original raw inputs archived for audit.
- Re-reporting when numbers change:
  - If only values change (new survey or new job posts): update metric value and publish new version with updated sample size and period; do not retroactively overwrite prior snapshots—keep historical pages.
  - If methodology changes: keep prior published numbers as “legacy” and present the new numbers separately; provide a reconciliation table showing the effect of the methodology change on key metrics.
- Minimum cell size rule: do not publish any subgroup statistic where n < 30; when 30 ≤ n < 100 publish with a “small sample” flag and wide confidence interval.

6) Statistics to explicitly reject (and why)
- Total industry revenue / market size estimates for wedding photography: reject. You have no platform booking data or tax/industry revenue data; any attempt would be a guess and not an original, reproducible measurement.
- Average profit margin across photographers: reject unless you collect business expense and revenue as part of survey (sensitive, low response and self‑reported). If you want to add it later, do it as binned revenue & expense categories with explicit opt-in and a clear privacy plan.
- Aggregations of other sites’ published numbers (e.g., average per‑wedding spend from a third party): reject as “aggregation” is not original and undermines the unique value of your page.
- “National average day rate” that mixes posted job-board rates and self-reported pay without separating sources—reject mixing sources into a single number. Present job-board posting rates and self-reported pay separately and label both clearly.

7) Presentation and citation practices to maximize citations
- For each metric headline show (1) the number, (2) n, (3) period, and (4) a one‑line method note (e.g., “Median of 3,124 job‑board postings with numeric day rates, 2022–2025”).
- Provide a “Methodology” button that opens the full methodology note (the exact sample, exclusion rules, question wording).
- Publish an open CSV/JSON for each snapshot with column-level metadata so journalists and researchers can re-use your exact published numbers.
- Encourage citation by providing a one‑sentence suggested citation (e.g., “Source: Site Name, Second‑Shooter Rate Snapshot v2026.08.01, job‑board n=3,124”).

8) Minimum sample reporting thresholds & statistical notes
- For proportions from the annual survey (n ~1,200–1,800): compute 95% binomial confidence intervals and publish them. Report margins of error for main proportions using standard formulas.
- For medians/percentiles, publish bootstrapped 95% intervals where useful (for sample sizes <300).
- When reporting regional job-board medians, publish sample size per region and flag cells with n<30.

9) Operational checklist to implement in March
- Add the exact new questions (G–K) to the March survey with forced numeric where specified.
- Implement required data validation on the survey (range checks).
- Set up a monthly job-board ETL that dedupes and standardizes day-rate fields and appends to a jobs database.
- Build a static “Statistics” page template that includes headline numbers, one-line method notes, and link to full methodology and CSV.
- Prepare the versioning/audit page and CSV export scripts.
- Decide whether to inflation-adjust older job-board rates (if so, define the index and methodology). If you do inflation adjustments, publish both nominal and inflation-adjusted numbers and document the index and base year.

Summary — quick list of publishable, original stats to start with
- Second-shooter day-rate distribution by US region (job-board; monthly updates; quarterly snapshots).
- Median and percentile distribution of main-shooter average booking price (annual survey; 5‑year trend).
- % offering packages vs a la carte (annual survey).
- Weddings shot per year (distribution; annual).
- Median editing hours per wedding (annual).
- % carrying business insurance (annual).
- Typical deposit percentage (add to March survey; annual).
- Typical gallery delivery time in days (add; annual).
- Frequency of hiring second shooters and pay cadence (add; annual).
- % outsourcing editing and median cost paid (add; annual).
- Price-increase cadence & typical percent raise (add; annual).

If you want, I can:
- Draft the exact JSON/CSV schema and column names for your export files so journalists can easily import.
- Draft the exact new survey question text and validation rules you can drop into your March survey form.
- Mock up the one‑page Methodology template for each metric (so the copy is ready to paste).

Which of the follow‑ups above would you like me to prepare next?

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Refresh brief for a page losing traffic

This page used to earn traffic and is fading. Write the refresh brief: which facts have gone stale and what current information replaces them; which sections the queries it still gets impressions for actually want, versus sections nobody asked for; what the page must add to beat whoever took its position — grounded in the competitor evidence provided, not guessed. Keep the URL and everything that still works; list what survives untouched so the refresh does not become a rewrite. If the evidence says the topic itself died, say that instead — a perfect refresh of a dead query recovers nothing. Page, stats, and competitor evidence: {{evidence}}

Fill in: The page's content summary, its query/impression history, and what currently outranks it.

Known limits: Distinguishing topic decay from ranking loss needs the query data — without impressions by query, the brief guesses which.

Run evidence

Sandbox run 2026-08-18 · openai/gpt-5-mini

Input fixture
PAGE: /guides/best-budget-3d-printers — last updated Nov 2024. Reviews
5 printers: Ender 3 V3 SE, Anycubic Kobra 2, Bambu A1 Mini, Prusa Mini+, Sovol SV06.
Includes a comparison table, first-print photos, and per-printer noise measurements.

STATS: Peaked at 9k impressions / 210 clicks per 28d in early 2025. Now 2.1k
impressions / 18 clicks. Still gets impressions for: "best budget 3d printer" (pos 14,
was 6), "bambu a1 mini review" (pos 8, stable), "ender 3 v3 se vs kobra 2" (pos 5,
stable). Lost entirely: "best 3d printer under 300".

COMPETITORS NOW ABOVE US for the head term: two 2026-dated roundups that include the
Bambu A1 (non-mini), Creality K1C, and Elegoo Neptune 4 Pro — none of which our page
mentions. Both show print-quality test photos on standardized models and cite print
speeds our page does not measure.

Model output (unedited)

Summary verdict
- The topic is alive — search interest remains but the page has been leapfrogged because competitors published fresh 2026 roundups that: 1) include newer, higher‑interest models (Bambu A1 [non‑mini], Creality K1C, Elegoo Neptune 4 Pro) that our page does not; and 2) publish standardized print-quality photos and print‑speed measurements that our page currently lacks. A refresh focused on adding those three elements (new contenders + standardized test photos + speed metrics), plus a short “best under $300” section, is required to recover impressions for the head term.
- Do not rewrite the whole URL or toss the existing content — much of it can survive. Below is a concrete, evidence‑grounded refresh brief: what’s stale, what must be added, what the queries actually want vs what’s unnecessary, and a prioritized task list.

1) Facts that have gone stale and the replacements required
- “Current” models and roundup freshness (stale)
  - Why stale: page last updated Nov 2024; competitors show 2026-dated roundups and now outrank us.
  - Fix: update page date + intro to state “updated 2026” and explicitly include the newer contenders named by competitors: Bambu A1 (non‑mini), Creality K1C, Elegoo Neptune 4 Pro. Add short bullets per new model with the same review template used for existing printers.

- Coverage gap: missing high-interest models (stale)
  - Why stale: both competing roundups include the Bambu A1 (non‑mini), Creality K1C and Elegoo Neptune 4 Pro; our page does not mention them.
  - Fix: add those three as review entries (or at minimum as a comparison-row and “Also consider” mini-reviews) and test them with the same protocol used for the current five printers.

- Missing print-speed data (stale)
  - Why stale: competitors explicitly cite print speeds and show standardized print photos; our page “does not measure” speed (per evidence).
  - Fix: add measured print-speed metrics and a short “print time for standard model at X quality settings” for each printer.

- Missing standardized print-quality photos (stale)
  - Why stale: competitors show standardized test-model photos; our first-print photos are not enough to compare across sites.
  - Fix: add standardized test prints (same model(s) across all printers) and side-by-side crops in the comparison table and in each review.

- “Best under $300” visibility (stale/omitted)
  - Why stale: we’ve lost entirely for that query. Likely our page does not include a clear under‑$300 section.
  - Fix: add a dedicated “Best 3D Printers under $300” subsection with tested contenders (research required to select current models), or a clear “budget picks” mini‑table with date-stamped price band.

2) Which sections the queries still want vs sections nobody asked for
- Sections queries actually want (based on the impressions we still get and competitor evidence)
  - Clear, dated roundup (2026) — people expect current-year guidance.
  - Direct comparisons between the headliners (example: Ender 3 V3 SE vs Anycubic Kobra 2) — our stable query for that comparison shows users want side‑by‑side specs + tests.
  - Dedicated, in-depth review for “Bambu A1 Mini” (query stable) — our Bambu A1 Mini content must be complete and accurate.
  - Measured print speed, standardized print-quality photos, and tradeoff commentary (speed vs quality).
  - A short “best for” recommendation labels (best for beginners, best for speed, best for silent operation, best under $300) so users can quickly pick.

- Sections nobody/rarely asked for (candidates to trim or move down)
  - Long historical brand background or multi‑paragraph company histories — unlikely to influence a “best budget” buying decision.
  - Very long assembly-only walkthroughs unconnected to objective performance — useful, but should be secondary or linked to separate build guides.
  - First-print photos alone as the sole quality evidence — these are less comparable than standardized test prints and are not what competitors show.

3) Exactly what the page must add to beat current competitors (grounded in competitor evidence)
From the competitor evidence we have, they succeeded by (A) adding Bambu A1, Creality K1C, Elegoo Neptune 4 Pro, and (B) providing standardized print photos and print-speed claims. To beat them we must match and surpass those elements:

- Add the missing contenders (must-do)
  - Add review entries (full template) for: Bambu A1 (non‑mini), Creality K1C, Elegoo Neptune 4 Pro.
  - For each, include the same set of specs you use for the five legacy printers so they’re directly comparable in the table.

- Measure and publish standardized print-quality photos (must-do)
  - Run the same standardized test models on every printer on the page (at minimum: one calibration/test model that shows overhangs, bridging, and fine detail; and a Benchy for baseline).
  - Display side‑by‑side high‑resolution crops for key features (stringing, overhang, bridging, surface finish).
  - Caption images with print settings used (layer height, temperature, filament brand/color, nozzle size).

- Measure and publish practical print-speed metrics (must-do)
  - For each printer, record: recommended max print speed, and measured print time for the standardized model(s) at two quality presets (e.g., “quality” = 0.12–0.16 mm, “fast” = 0.24–0.28 mm), using the same slicer/profile across printers.
  - Add a “speed vs. quality” short chart or table entry.

- Expand the comparison table (must-do)
  - Add columns: Measured print-speed (time for test model), Standardized print-quality score (see methodology below), Date-tested, Price band (with date), and Keep per-printer noise (unique asset).
  - Make a sortable table (speed, quality, price) so users can quickly filter.

- Add a short, transparent methodology box (must-do)
  - State test models, filament, slicer and profile, layer heights tested, ambient conditions, and how the quality score is computed (objective criteria + pass/fail checklist).
  - This matches competitors who present standardized tests and gives credibility.

- Add a “Best under $300” budget section (must-do)
  - Research current sub‑$300 contenders and test 1–3. If you can’t test immediately, temporarily add a “shortlist” POSITIVE/NEGATIVE summary and mark as “To be tested — updated 2026”.
  - Make sure the price band is date-stamped (e.g., “As of Aug 2026”).

- Keep and highlight unique assets we already have (must-do)
  - Keep the per-printer noise measurements (unique differentiator). Make noise data more visible in the comparison table and call these out in “best for quiet” recommendations.
  - Keep and reuse existing first-print photos as “first impressions” alongside the new standardized test photos.

- On-page UX and SEO tweaks (must-do)
  - Update title/URL meta & H1 to include “2026” or “Updated Aug 2026” and add an easily scannable TL;DR shortlist at top with “best overall / best budget / best quiet / best for speed”.
  - Add schema.org review/product structured data including test dates and measured metrics.
  - Internally link to detailed printer pages (if they exist) and to budget-specific content (to recapture the under-$300 query).

4) What survives untouched (don’t throw away)
- URL: /guides/best-budget-3d-printers — keep this canonical URL.
- The five existing reviews (Ender 3 V3 SE, Anycubic Kobra 2, Bambu A1 Mini, Prusa Mini+, Sovol SV06) — keep copy and headline; update them only where needed to add new metrics and photos.
- Existing comparison table structure — keep, but expand columns/data (don’t replace structure).
- Existing first-print photos — keep as supplementary “first-print” evidence.
- Existing per-printer noise measurements — preserve intact and surface them in the revised comparison table and “best for silence” callout.
- Any internal links and backlinks already pointing to the page — preserve (do not change canonical or URL).

5) What can be removed, moved, or demoted
- Move long assembly walkthroughs to separate “how to assemble” posts and link out, or collapse into an accordion — keep on-page focus on buying-relevant performance metrics.
- Demote long brand histories or marketing copy — shorten to one-line context sentences.

6) Prioritized refresh task list (quick sprint)
Priority 1 (high impact, 3–7 working days)
- Add Bambu A1 (non‑mini), Creality K1C, Elegoo Neptune 4 Pro entries (copy + spec rows).
- Run standardized tests on all existing printers + the three new ones: print standardized model(s), capture photos.
- Measure practical print times for the standardized model(s) at two presets.
- Update comparison table with new columns (speed, standardized quality score, price band, tested date).
- Add methodology box explaining test protocols.

Priority 2 (medium impact, 2–4 working days)
- Add “Best under $300” section (test one or two models; if testing not possible immediately, add short, date-stamped shortlist).
- Add sortable/filterable table features and “best for” TL;DR at top.
- Surface noise measurements as a distinct “best quiet” recommendation.

Priority 3 (lower impact, 1–3 working days)
- Update meta title, description, H1 to include “Updated 2026”.
- Add structured data (review/product schema including metrics).
- Move or collapse assembly/brand history content.

7) Quick copy/UX changes to improve CTR (do at publish)
- Update top-line TL;DR: one-line winner + 1-line budget pick + 1-line quiet pick + “Updated Aug 2026”.
- Add badges in table for “Best overall”, “Best budget”, “Best silent”, “Best speed” based on measured metrics.
- Caption standardized photos with exact print settings and date; use “view large” toggle to encourage clicks.

8) Measurement: how to judge refresh success
- Short term (2–6 weeks): recover impressions for “best budget 3d printer” toward previous range (goal: double current impressions) and hold/improve positions for “bambu a1 mini review” and “ender 3 v3 se vs kobra 2”.
- Medium term (6–12 weeks): reclaim a top‑10 position for the head term by matching competitor freshness and test evidence.
- Track clicks, CTR, impressions, and SERP positions for the three queries called out in the stats and for “best 3d printer under 300”.

9) If you can’t test immediately — temporary tactical steps
- Add a clearly labeled “To be tested (Aug 2026)” block for the three competitor models and the under-$300 picks, with short data pulled from manufacturer specs and price bands (date-stamped). This prevents empty gaps and signals freshness. Priority: still run the tests as soon as possible and replace provisional content.

Summary action checklist (minimum viable refresh to compete)
- Add Bambu A1 (non‑mini), Creality K1C, Elegoo Neptune 4 Pro entries.
- Run standardized print tests + capture photos for all printers listed.
- Measure and publish practical print-time metrics for a standardized model at two presets.
- Expand comparison table to include speed, quality score, tested date and price band.
- Add methodology box and update page date to 2026.
- Add or test and publish “Best under $300” picks.
- Keep URL, existing reviews, noise data, and first-print photos intact.

If you want, I can:
- Draft the updated TL;DR and H1/meta lines for 2026.
- Produce the methodology text (copy you can paste into the page).
- Create the new comparison table column layout and scoring rubric for the standardized print-quality score. Which of those would you like next?

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)