PromptZone - AI Prompts, Guides and Tools for Builders

Code review and debugging

Most prompt libraries treat coding as code generation. That is the part you least need help with — you can already write the function. The expensive hours go somewhere else: reading an unfamiliar file well enough to change it safely, working out why a test only fails in CI, deciding whether a migration will lock a table in production, and reviewing a diff carefully enough that the bug does not reach main.

These prompts are for those moments. Every one takes real input — a diff, a stack trace, a query plan, a failing test — and asks the model to do something specific with it. They are written to make the model say "I cannot tell from this" rather than invent a plausible cause, because a confident wrong answer about why production broke costs more than no answer.

Curated

Assess the blast radius before deleting code

You are deciding whether this code is safe to delete. List everything that could still depend on it, including callers in the code given, dynamic dispatch by name, serialised references such as enqueued jobs or stored class names, database columns or migrations, public API surface, and anything reachable from configuration or templates. For each, say how you would verify it before deleting. End with a plain verdict: safe, unsafe, or cannot determine from what was provided. Code: {{code}}

Fill in: The code you want to remove, plus how it is referenced if you already know of any callers.

Known limits: Cannot grep your repository — dynamic references are the ones it will raise as questions rather than facts.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
# Candidate for deletion: this worker and its class.

class LegacyDigestWorker
  include Sidekiq::Job
  sidekiq_options queue: "mailers", retry: 5

  def perform(user_id, digest_type = "weekly")
    user = User.find_by(id: user_id)
    return unless user

    DigestMailer.send(digest_type, user).deliver_now
  end
end

# Known: nothing in app/ calls LegacyDigestWorker.perform_async directly.
# The scheduler config has an entry keyed "legacy_digest" that was commented out
# three months ago. Sidekiq's retry set currently holds ~400 jobs.

Model output (unedited)

Summary of the code
- LegacyDigestWorker is a Sidekiq worker in the mailers queue.
- It takes a user_id and an optional digest_type (default "weekly"), finds the user, and sends a digest mail via DigestMailer.send(digest_type, user).deliver_now.

What could still depend on it and how to verify before deleting

1) In-code callers (static references)
- What to check
  - Any code that calls LegacyDigestWorker.perform_async (or any dynamic dispatch that ends up enqueuing this worker).
  - Any references to the class name as a string or symbol (e.g., "LegacyDigestWorker") anywhere that calls perform_async or performs a constantize.
- How to verify
  - Full repo search for:
    - perform_async calls: rg -n "LegacyDigestWorker\.perform_async" -S
    - Any reference to the class name string: rg -n "LegacyDigestWorker" -S
  - If found, inspect and determine if those code paths should be migrated or removed along with the worker.
- What if you find references
  - If legitimate enqueue points exist, you’ll need a migration plan (e.g., point to a new worker or remove those calls).

2) Dynamic dispatch by name / serialized job payloads in Sidekiq
- What to check
  - Sidekiq jobs already enqueued or retried (serialized in Redis) that refer to this class.
  - Specifically, items with klass = "LegacyDigestWorker" in RetrySet, ScheduledSet, DeadSet, or any live queues.
- How to verify
  - In Rails console or a script:
    - require 'sidekiq/api'
    - Sidekiq::RetrySet.new.each { |j| puts "retry #{j.jid} #{j.klass}" if j.klass == 'LegacyDigestWorker' }
    - Sidekiq::ScheduledSet.new.each { |j| puts "scheduled #{j.jid} #{j.klass}" if j.klass == 'LegacyDigestWorker' }
    - Sidekiq::DeadSet.new.each { |j| puts "dead #{j.jid} #{j.klass}" if j.klass == 'LegacyDigestWorker' }
    - Sidekiq::Queue.new.each { |j| puts "queue #{j.jid} #{j.klass}" if j.klass == 'LegacyDigestWorker' }
  - Alternatively, use the Sidekiq web UI to inspect the Retry, Scheduled, and Dead sets and filter by class.
- What if you find references
  - Do not delete yet. Decide whether to:
    - Let those jobs run (and complete or fail) before deleting, or
    - Requeue/migrate them to a new worker, or
    - Remove the enqueue points and then delete the worker.
  - If you must delete, you’ll need to purge or migrate those jobs first.

3) Scheduler configuration (configured/delayed runs)
- What to check
  - Any Sidekiq schedule entries that reference legacy_digest, even if commented out, or any code that loads a schedule file containing this key.
  - The presence of a sidekiq.yml or an initializer that loads a YAML/JSON schedule including a legacy_digest entry.
- How to verify
  - Grep for the key across the repo:
    - rg -n "legacy_digest" -S
  - Inspect config/sidekiq.yml, config/initializers/sidekiq.rb, or any sidekiq-cron/sidekiq-scheduler config files to see if legacy_digest would be loaded or scheduled.
  - If there is a scheduled job present in Redis, verify whether that schedule is still intended to run and if it should be migrated.
- What if you find references
  - If a schedule will re-create the job, you’ll need to remove or migrate that schedule first (and ensure the job won’t be enqueued after deletion).

4) Database migrations / columns (unlikely, but verify)
- What to check
  - Any migrations or schema that mention LegacyDigestWorker (very unlikely, but possible if the design stored references to workers).
- How to verify
  - Search migrations: rg -n "LegacyDigestWorker" db/migrate -S
- What if you find references
  - Review the intent; typically there won’t be a dependency on a worker class in the DB, but if found, adjust migrations or remove the dependency.

5) Public API surface / external usage
- What to check
  - Any public API, webhook, or external service that enqueues this worker by name or expects a digest to be sent on a schedule.
  - Any documentation, README, or API clients that reference the LegacyDigestWorker.
- How to verify
  - Search documentation and external-facing code for mentions of “legacy_digest” or “LegacyDigestWorker”.
  - If the app exposes a REST/GraphQL endpoint that triggers digests, verify whether it routes through a worker or a direct mailer call.
- What if you find references
  - Decide whether to migrate those calls or update external clients to use a different worker or approach.

6) Templates / configuration reachability
- What to check
  - Any templating or config that relies on the presence of a named job (e.g., templates showing digest links, or configuration that implies a digest is sent).
- How to verify
  - Scan templates for mentions of digest types or the worker class name.
  - Confirm there are no template-driven enqueues of this worker.
- What if you find references
  - If templates or config depend on this worker to exist, plan an alternate behavior or feature toggle before deletion.

7) Tests and test fixtures
- What to check
  - Specs that reference LegacyDigestWorker or its behavior (e.g., tests that expect a digest email to be enqueued or delivered via the worker).
- How to verify
  - Run test suite or search tests for "LegacyDigestWorker" or mailer digests in tests.
- What if you find references
  - Update tests or ensure they are updated to reflect removal of this worker.

8) Observability and maintenance hooks
- What to check
  - Any logging, metrics, or dashboards that track digest sending via this worker.
- How to verify
  - Review logging statements or metrics that mention LegacyDigestWorker or digest_type.
- What if you find references
  - Decide whether to migrate logs/metrics to a new target if you replace the worker.

Notes on the given facts
- The only caller information provided says: nothing in app/ calls LegacyDigestWorker.perform_async directly.
  - This reduces risk of static code references, but dynamic/job-serialized references still matter.
- The scheduler config had a legacy_digest entry that was commented out three months ago.
  - If that schedule is still loaded anywhere, it could re-enqueue the worker in the future; verify by inspecting scheduler config files and any runtime loading of schedules (e.g., sidekiq.yml, sidekiq-scheduler initializer).
- Sidekiq’s retry set currently holds ~400 jobs.
  - This is a key red flag: some of those jobs may be for LegacyDigestWorker. If any exist, deleting the worker would cause those jobs to fail when retried or require a migration of those jobs before deletion.

Plain verdict
- Cannot determine safely from what was provided.
- There is a real risk that there are enqueued/retry/scheduled jobs referencing LegacyDigestWorker (especially given the non-trivial retry set and a previously configured scheduler entry). Without verifying the contents of the Sidekiq queues and scheduled tasks, deleting the worker could cause deserialization errors or orphaned jobs.

Recommended next steps (before deleting)
- Inspect repo for references to perform_async and the class name (as described above).
- Inspect Sidekiq queues, retry, scheduled, and dead sets for any jobs with klass == "LegacyDigestWorker".
- Confirm the status of any sidekiq-scheduler or sidekiq.yml entries related to legacy_digest; remove or migrate if necessary.
- If there are any dependent jobs or references, address them ( migrate to a new worker, remove enqueue points, or purge the jobs) before deletion.
- Only after confirming no references exist (or after migrating them) should you delete the worker.

Final answer: Cannot determine.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Check a PR description against its diff

Compare this pull request description against what the diff actually does. Report three things: changes present in the diff but not mentioned in the description, claims in the description not supported by the diff, and any change that is larger in blast radius than the description implies. Quote the specific lines for each. If the description accurately covers the diff, say so. PR: {{pr}}

Fill in: The description and the full diff.

Known limits: A description can be accurate and the change still wrong; this checks correspondence, not correctness.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
DESCRIPTION: "Fix typo in the welcome email subject line."

--- a/app/mailers/welcome_mailer.rb
+++ b/app/mailers/welcome_mailer.rb
@@ -3,7 +3,7 @@ class WelcomeMailer < ApplicationMailer
   def welcome(user)
     @user = user
-    mail(to: @user.email, subject: "Welcom to PromptZone")
+    mail(to: @user.email, subject: "Welcome to PromptZone")
   end
 end
--- a/config/initializers/mail.rb
+++ b/config/initializers/mail.rb
@@ -8,5 +8,5 @@
-  config.action_mailer.perform_deliveries = true
+  config.action_mailer.perform_deliveries = Rails.env.production?
--- a/app/models/user.rb
+++ b/app/models/user.rb
@@ -41,7 +41,6 @@ class User < ApplicationRecord
-  validates :email, uniqueness: { case_sensitive: false }

Model output (unedited)

Here’s a focused comparison of the PR description against the actual diff.

1) Changes present in the diff but not mentioned in the description
- app/mailers/welcome_mailer.rb
  - -    mail(to: @user.email, subject: "Welcom to PromptZone")
  - +    mail(to: @user.email, subject: "Welcome to PromptZone")
  Quote lines from diff:
  - "mail(to: @user.email, subject: "Welcom to PromptZone")"
  - "mail(to: @user.email, subject: "Welcome to PromptZone")"

- config/initializers/mail.rb
  - -  config.action_mailer.perform_deliveries = true
  - +  config.action_mailer.perform_deliveries = Rails.env.production?
  Quote lines from diff:
  - "config.action_mailer.perform_deliveries = true"
  - "config.action_mailer.perform_deliveries = Rails.env.production?"

- app/models/user.rb
  - -  validates :email, uniqueness: { case_sensitive: false }
  Quote line from diff:
  - "validates :email, uniqueness: { case_sensitive: false }"

What these mean:
- In addition to fixing the typo in the subject, the diff introduces a production-only mail delivery setting and removes a uniqueness validation on the User email. These are not mentioned in the description.

2) Claims in the description not supported by the diff
- None. The description claims only a typo fix for the welcome email subject. The diff indeed fixes that piece, but it also includes additional, unmentioned changes (as shown above). If you only consider the explicit claim, it is supported; however the description does not claim anything about the other changes.

3) Any change that is larger in blast radius than the description implies
- The initializer change in config/initializers/mail.rb
  - "-  config.action_mailer.perform_deliveries = true"
  - "+  config.action_mailer.perform_deliveries = Rails.env.production?"
  This changes behavior (delivery is only enabled in production), affecting mail delivery in environments beyond just the subject string.

- The removal of email uniqueness validation in app/models/user.rb
  - "-  validates :email, uniqueness: { case_sensitive: false }"
  Removing a validation has broad implications for user data integrity across the app (e.g., permitting duplicate emails unless a separate constraint exists).

- Given description only mentions fixing a typo, these two changes represent a larger blast radius than the description implies.

4) Overall assessment
- The description partially matches the diff (the subject typo is corrected), but it omits two other substantial changes. Therefore, the description does not fully cover the diff. The subject fix is accurate, but the PR also makes wider changes that the description does not mention.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Explain a file you did not write

You are explaining this file to an engineer who has to modify it today and has never seen it. In this order: one sentence on what it is responsible for; the public entry points and who calls them; the state it reads and the state it mutates; the non-obvious behaviour or assumption a newcomer would most likely break; and the one thing you would verify before changing it. Mark anything you are inferring from naming rather than reading as an inference. Code: {{code}}

Fill in: The file, ideally whole rather than excerpted.

Known limits: Infers caller behaviour it cannot see; treat the 'who calls this' section as a hypothesis to check, not a finding.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
class SessionRefresher
  REFRESH_WINDOW = 5.minutes

  def initialize(session, clock: Time)
    @session = session
    @clock = clock
  end

  def call
    return :expired if @session.expires_at < @clock.now
    return :fresh if @session.expires_at - @clock.now > REFRESH_WINDOW

    @session.with_lock do
      @session.update!(expires_at: @clock.now + 1.hour, refreshed_count: @session.refreshed_count + 1)
    end
    :refreshed
  end
end

Model output (unedited)

- What it is responsible for (one sentence): SessionRefresher is a small service object that refreshes a session’s expiration within a short window and reports the outcome. (inference from naming)

- Public entry points and who calls them:
  - Public entry point: the instance method call (SessionRefresher#call).
  - Who calls them: not shown in this file; typically invoked by higher-level code that constructs a SessionRefresher with a session (inference based on usage pattern).

- State it reads and the state it mutates:
  - Reads:
    - @session.expires_at
    - @clock.now
  - Mutates:
    - @session.expires_at (set to @clock.now + 1.hour)
    - @session.refreshed_count (incremented by 1)
  - Also uses @session.with_lock to perform the update atomically via update!

- The non-obvious behaviour or assumption a newcomer would most likely break:
  - It uses three branches: :expired when expires_at < now; :fresh when expires_at - now > 5 minutes; otherwise it refreshes by extending expires_at to now + 1 hour and incrementing refreshed_count, then returns :refreshed. A notable detail is that if expires_at == now, it is treated as not expired and will be refreshed (boundary condition). It relies on a DB-level lock around the update to ensure atomicity.

- The one thing you would verify before changing it:
  - Verify the boundary semantics around expiration (specifically whether expires_at == now should be considered expired or not) and ensure that the 5-minute window and 1-hour extension align with current business rules (this is the most likely source of subtle bugs).

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Explain what a regex actually matches

Explain this regular expression in plain language, then give a table of five inputs it matches and five near-miss inputs it does not, choosing the near-misses to expose its edge cases rather than obvious non-matches. Flag any catastrophic backtracking risk with the input shape that triggers it, and note anything that behaves differently across regex engines. Regex and its intended purpose: {{regex}}

Fill in: The expression and what it is meant to match — the gap between those two is usually the bug.

Known limits: Engine differences are real; verify anything version-specific against your actual runtime.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
^(?:[a-zA-Z0-9!#$%&'*+/=?^_`{|}~-]+(?:\.[a-zA-Z0-9!#$%&'*+/=?^_`{|}~-]+)*)@(?:[a-zA-Z0-9](?:[a-zA-Z0-9-]*[a-zA-Z0-9])?\.)+[a-zA-Z]{2,}$

Intended purpose: validating email addresses on a signup form.

Model output (unedited)

Plain-language explanation

- Overall: The pattern checks a typical email form: a local part, an @, and a domain with a valid-looking host and a real-sounding top-level domain (TLD).

- Local part (before @):
  - One or more characters from a fairly wide ASCII set (letters, digits, and many punctuation characters).
  - Dots are allowed, but only as separators between dot-separated segments; there cannot be consecutive dots or a leading/trailing dot in the local part.
  - In other words, the local part is made of one or more dot-separated chunks, each chunk consisting of allowed characters.

- @ symbol: A literal at-sign separates local part and domain.

- Domain part (after @):
  - One or more labels separated by dots. Each label:
    - starts with an alphanumeric character,
    - may contain alphanumeric characters or hyphens inside,
    - cannot end with a hyphen.
  - There must be at least one dot in the domain (so something like example.com, not just example).
  - The final part (the TLD) must be at least two ASCII letters (no digits or hyphens in the TLD).

- Summary of what it accepts:
  - Local part with typical characters and dot-separated segments.
  - Domain with one or more dot-separated labels, no trailing hyphen in a label.
  - TLD that is at least two letters.

- Summary of what it forbids (highlights edge behavior):
  - No leading/trailing dot in local part; no double dots in the local part.
  - No domain label starting or ending with a hyphen.
  - No digits or hyphens in the TLD (must be letters only, at least 2).
  - No domain without at least one dot (e.g., something@localhost is rejected).

Five inputs it matches (examples)

- [email protected]
- [email protected]
- customer/[email protected]
- [email protected]
- [email protected]

Five near-misses (inputs that fail, showing edge cases)

- [email protected] (local part starts with a dot)
- [email protected] (local part ends with a dot)
- [email protected] (domain label ends with a hyphen)
- [email protected] (TLD too short; only 1 letter)
- [email protected] (TLD contains digits)

Notes on edge cases and real-world subtleties

- This regex is a pragmatic email validator for signup forms, but not a perfect RFC 5322 validator. It rejects some valid-but-uncommon cases (e.g., quoted strings in the local part) in favor of a simpler, safer rule set.

- It enforces a simple, ASCII-only domain and TLD. It will not match internationalized domain names (IDNs) or punycode unless they are converted to ASCII (punycode) first.

- It allows many characters in the local part (including some you’ll rarely see in practice, like certain punctuation), which some real-world mail servers might reject even though this pattern accepts them.

- It requires a TLD of at least two letters. That excludes some rare or old TLDs, and also excludes numeric TLDs.

- It supports multi-part TLDs (like co.uk) by allowing multiple dot-separated domain labels, with a final letter-only TLD.

Catastrophic backtracking risk (engineer’s caution)

- The pattern is backtracking-friendly (i.e., it uses a lot of backtracking to try alternative segmentations) and does not use atomic groups or possessive quantifiers to pin down matching as aggressively as possible.

- In theory, extremely long inputs crafted to maximize ambiguity can trigger significant backtracking in some regex engines (JavaScript, Python’s re, PCRE, .NET) and cause performance degradation.

- A shape that can illustrate the risk (shape only, not a single compact string you’d type): a very long local part consisting of many identical characters with many dot-separated segments (for example, a.a.a.a....a) and a similarly long domain with many dot-separated segments, ending with a non-letter TLD (e.g., digits) so the match almost succeeds but fails late. The engine may try many different ways to partition the local and domain segments before concluding a mismatch, potentially exploding runtime as input length grows.

  - Conceptual example shape: a long local-part of repeated "a" separated by dots (a.a.a....a) + @ + a long domain of repeated "b" separated by dots (b.b.b....b.123) where the final TLD "123" is invalid. The more segments you add, the more backtracking paths the engine explores.

  - Practical note: If you expect untrusted inputs or have very large inputs, consider adding lightweight pre-checks (e.g., quick length checks on local part and domain) or using a non-backtracking validator or a well-maintained email-validation library to avoid pathological cases.

Engine differences to be aware of

- Case handling: This pattern explicitly includes both uppercase and lowercase letters in the character classes, so it is effectively case-insensitive with respect to the allowed letters without relying on a case-insensitive flag.

- Unicode handling: The regex uses only ASCII ranges. If you enable Unicode mode (various engines) and expect internationalized inputs, you’ll need a different approach or IDN handling first. In most engines, the character classes here remain ASCII-only regardless of Unicode mode.

- Performance and backtracking: JavaScript, Python, PCRE, and .NET engines differ in how aggressively they backtrack. The structure here is not atomic, so some engines may exhibit more backtracking than others on crafted inputs. If you hit performance issues, consider refactoring with atomic groups or using a dedicated email validator.

- Compiler differences: The literal dash inside the first character class is placed at the end, which keeps it as a literal dash rather than indicating a range. Most engines treat it correctly, but when editing patterns, ensure the dash isn’t accidentally interpreted as a range (especially if you move characters around inside that class).

- RFC compliance vs. practical use: RFC 5322 allows very broad local-parts and domain syntax, but this pattern opts for a safer, more common-sense subset suitable for signup forms. If you need to support quoted strings in the local part or IDNs, you’ll need a different approach.

Short recommendations

- For signup forms, this regex is a reasonable balance of practicality and safety, but consider:
  - adding a length limit on the local part (<= 64) and domain (overall <= 255) as a first filter,
  - and/or using a reputable, well-tested email-validation library that implements modern standards and uses non-backtracking checks where possible,
  - and possibly handling IDNs (punycode) before validation if you intend to support international email addresses.

If you want, I can provide a shorter, more permissive or more strict variant (for example, a version with a non-backtracking approach or one that supports IDNs).

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Explain why a test is flaky

Act as an engineer who has seen every common source of flakiness: shared mutable state between examples, real clock or timezone dependence, ordering assumptions on unordered results, unawaited async work, network or filesystem access, and random data. Given this test and what is known about when it fails, rank the likely causes and give, for each, the specific line that would produce it and a deterministic rewrite. If the failure pattern does not match any of these, say so. Test and failure pattern: {{test}}

Fill in: The test source plus when it fails — CI only, in parallel, after a particular other test, or at certain times of day.

Known limits: Without the failure pattern it lists generic causes; the 'when it fails' half of the input is what makes the answer specific.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
it "lists the three most recent posts" do
  create(:post, title: "old")
  create(:post, title: "middle")
  create(:post, title: "new")

  get "/api/posts/recent"

  expect(json_body.map { |p| p["title"] }).to eq(%w[new middle old])
end

# Passes locally every time. Fails in CI roughly 1 run in 8, always with
# "middle" and "new" swapped. CI runs specs in parallel across 4 processes.
# Post has `default_scope { order(created_at: :desc) }`.

Model output (unedited)

Short answer: the failure pattern matches ordering assumptions on unordered results caused by a non-deterministic sort. The most likely root cause is the model’s default_scope ordering only by created_at, which can produce ties in CI (especially with parallel tests) and leads to swapped positions for posts with identical timestamps.

Ranked causes (with line(s) that would produce the issue and a deterministic rewrite)

1) Ordering assumptions on unordered results (most likely)
- Why it matches: In CI, inserts happen in parallel. If three posts are created within the same clock resolution (created_at), the database may return rows with equal created_at in any order. The default_scope orders by created_at only, so the relative order of ties is undefined and can flip across runs (e.g., swapping “middle” and “new”).
- The line that produces it:
  - app/models/post.rb: default_scope { order(created_at: :desc) }
- A deterministic rewrite (two equivalent options):
  - Option A: Use a stable tie-breaker in the default order
    - app/models/post.rb
      - default_scope { order(created_at: :desc, id: :desc) }
    - Rationale: adding id as a secondary key makes the sort deterministic when created_at ties occur.
  - Option B: Avoid relying on a global default scope and order explicitly in queries
    - app/models/post.rb
      - remove or replace the default_scope with a explicit scope:
        - scope :recent, -> { order(created_at: :desc, id: :desc) }
    - In controller/API (or in the test’s target query):
      - Post.recent.limit(3)  # or whatever the actual query is
- Optional test-side improvement (not required but helps): test against a set or explicitly compare sequences with a deterministic order, or test content without assuming exact order.

2) Shared mutable state between examples
- Why it doesn’t apply here: The failure is a consistent alternation in order observed across CI runs, not a hidden mutation leaking between example blocks in the same process. Rails tests typically rely on transactional tests or DB cleanup to isolate examples; there’s no obvious shared in-memory state here causing the swap of two items.
- The line that would produce it (if it were the cause): not applicable to this pattern in this particular code.
- Deterministic rewrite (if you want to rule it out): ensure proper test isolation (use transactional tests in each process; ensure per-process DB cleanup).

3) Real clock or timezone dependence
- Why it doesn’t apply: The test depends on ordering by created_at, but the failure is not due to time zone conversions or clock drift; it’s simply a non-deterministic ordering when timestamps tie.
- The line that would produce it: not applicable.
- Deterministic rewrite: ensure deterministic ordering as in (1); timezone handling would only matter if you were formatting times or filtering by a time window, which isn’t the case here.

4) Unawaited async work
- Why it doesn’t apply: This is a synchronous API test; there’s no asynchronous background work in the test path shown.
- The line that would produce it: not applicable.
- Deterministic rewrite: if you had async side-effects, you’d await or synchronize them; not relevant here.

5) Network or filesystem access
- Why it doesn’t apply: The test is purely database-backed and HTTP API-based; no network or FS IO is involved in the failure described.
- The line that would produce it: not applicable.
- Deterministic rewrite: remove any external IO from the path; not applicable.

6) Random data
- Why it doesn’t apply: The test uses fixed creates; there’s no random data generation affecting the order here.
- The line that would produce it: not applicable.
- Deterministic rewrite: make data deterministic (seeded RNG, or remove randomness). Not relevant.

What to do next
- Implement the deterministic ordering fix (prefer a stable tie-breaker). For example:
  - Change: default_scope { order(created_at: :desc) }
  - To: default_scope { order(created_at: :desc, id: :desc) }
  - Or remove the default_scope and use an explicit scope in queries: scope :recent, -> { order(created_at: :desc, id: :desc) } and call Post.recent.limit(3) in the API.
- Optional: adjust tests to assert on the exact content order only when you’ve guaranteed a deterministic order, or switch to asserting the presence of the expected set of titles rather than their exact sequence.

If you want, I can suggest concrete edits for the model, the controller/endpoint, and the spec to lock this down end-to-end.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Find the root cause from a stack trace

Act as an engineer on call. Given this stack trace, identify the most likely root cause and rank your top three candidates by probability. For each: name the specific frame that supports it, state what you would check to confirm or eliminate it, and say what evidence in the trace would have to be different for it to be wrong. Do not propose a fix until the cause is established. If the trace is insufficient, say what is missing. Trace: {{trace}}

Fill in: The full trace including the exception class and message, not just the top frames.

Known limits: Truncated traces produce confident guesses; paste the whole thing, including any 'caused by' chain.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
ActiveRecord::StatementInvalid: PG::TRDeadlockDetected: ERROR:  deadlock detected
DETAIL:  Process 4821 waits for ShareLock on transaction 99182; blocked by process 4930.
Process 4930 waits for ShareLock on transaction 99180; blocked by process 4821.
  app/services/inventory/reserve.rb:34:in `block in call'
  app/services/inventory/reserve.rb:31:in `each'
  app/services/inventory/reserve.rb:31:in `call'
  app/controllers/checkouts_controller.rb:22:in `create'
Occurs roughly 12 times/hour, only between 18:00-20:00 UTC, only on the checkout path.
reserve.rb:31 iterates cart line items and calls `item.with_lock { item.decrement!(:stock) }`.

Model output (unedited)

Short answer
- Most likely root cause: Deadlock caused by reserving multiple inventory items in a single checkout, where each item is locked with with_lock and the two concurrent checkouts lock items in different orders (classic multi-row deadlock). The stack shows the lock boundary clearly in reserve.rb and the deadlock detail shows two processes waiting on ShareLocks in opposite directions.
- Other plausible candidates are variations of the same root (lock-order rotations on multiple items, or longer transactions holding locks). There isn’t enough trace data to decisively rule them all out yet.

Top candidates (ranked by probability)

1) Multi-item lock order deadlock during inventory.reserve
- Why this fits the trace
  - The code path explicitly locks each item in the cart: reserve.rb lines 31-34 call item.with_lock { item.decrement!(:stock) } inside a loop over cart line items.
  - Deadlocks arise when two concurrent checkouts each lock two (or more) items but in different orders, forming a cycle (A locks, waits for B; B locks, waits for A).
  - The trace shows two processes waiting on ShareLock on different transactions, which is the classic symptom of cross-row locking in a loop like this.
- What I would check to confirm/eliminate
  - Confirm that carts in the failing window contain multiple line items. If failures occur only for single-item carts, this weakens the hypothesis.
  - Inspect PostgreSQL deadlock logs for the exact item_ids involved and the lock acquisition order (which items were locked first by each transaction, and which item each is waiting on).
  - Look for a consistent ordering in the lock acquisition across failures. If two concurrent requests lock items in the same order, deadlocks should be rare or absent.
  - Reproduce in a test: simulate two concurrent checkout requests where each cart contains multiple items, with items ordered differently, and observe if a deadlock occurs.
- What evidence would need to be different to disprove
  - If all deadlocks involve only a single item per checkout (no cross-item locking), or if the deadlock cycles never involve more than one item, the root cause would be different.
  - If the deadlocks occur even when every checkout locks items in the same deterministic order (e.g., always by item_id ascending) the current hypothesis would be weakened.
- What’s missing in the trace
  - The specific item_ids involved in the locks.
  - The exact lock order and the SQL emitted (e.g., SELECT ... FOR UPDATE or FOR SHARE) for each item.
  - The outer transaction boundary around reserve (is there a surrounding transaction that keeps locks open longer?).

2) Deadlock between two transactions locking two items in opposite order (A then B vs B then A)
- Why this fits the trace
  - Very close variant of #1: two different items are locked, and the cycle is created by reversed lock order between two concurrent checkouts.
  - The frame reserve.rb:31-34 strongly supports that multiple item locks can happen in one call.
- What I would check to confirm/eliminate
  - Same checks as #1: item_ids involved, order of locking per transaction, and exact deadlock graph from pg_logs.
  - Verify whether the cycle always contains exactly two items or sometimes more; cycles with more items still fit but provide more constraint on root cause.
- What evidence would need to be different to disprove
  - If every deadlock involves more than two distinct items and the cycle never reduces to A<->B, or if the logs show a deadlock with a single item and a different resource (not item rows), this would point to a different root cause.
- What’s missing in the trace
  - Item-level lock details (ids) and the sequence of lock acquisitions.

3) Long-running/outer-transaction issue causing or exacerbating deadlocks
- Why this could be relevant
  - The failure occurs only on checkout and during a narrow time window, suggesting there may be bursts of concurrency or long transactions that hold locks longer, increasing chances of deadlock.
  - If there is an outer transaction around the entire reserve flow, multiple item.with_lock blocks could be nested inside a larger transaction, complicating lock ordering and increasing deadlock risk.
- What I would check to confirm/eliminate
  - Look at the controller action (checkouts_controller#create) to see if it wraps the reserve call in a larger transaction that spans multiple items and perhaps other tables.
  - Check for any long-running transactions on pg_stat_activity during the failure window.
  - Confirm whether with_lock calls are truly isolated per item (i.e., each item lock is its own short transaction) or whether an outer transaction keeps locks held longer.
- What evidence would need to be different to disprove
  - If deadlocks occur even when the reserve calls are clearly wrapped in a small, isolated transaction (or not wrapped at all by a larger transaction), this weakens this hypothesis.
  - If there are no long-running transactions visible in pg_stat_activity during failures, this would also weaken it.
- What’s missing in the trace
  - The controller’s transaction boundaries and any surrounding transaction blocks.
  - Real-time pg_stat_activity or pg_locks samples showing long-running transactions.

What’s missing in the trace (overall)
- Exact item_ids involved in the deadlocks.
- The sequence/order of locks acquired by each process (which item_id each process locked first).
- The SQL emitted by item.with_lock (the actual FOR UPDATE/FOR SHARE statements).
- Whether there is an outer transaction wrapping the entire reserve call (or multiple item locks) in checkouts_controller#create.
- PostgreSQL deadlock graph from the logs (or pg_locks/pg_stat_activity context) to clearly map the cycle.

What to do next (to confirm the root cause)
- Enable richer deadlock logging if not already on:
  - Turn on log_lock_waits and log_min_duration_statement in PostgreSQL, and capture the deadlock graph in the PostgreSQL logs.
  - Correlate the failing timestamps with the SQL statements issued for each item in the cart (which item_ids, and in what order).
- Instrumentation in code (non-breaking, for debugging only):
  - Add minimal logging around the block in reserve.rb to log the item_id being locked and the order, for each checkout attempt.
  - Consider emitting a one-time warning if a cart has more than one line item to help quantify how often multi-item carts are involved in failures.
- Short-term mitigation (to gather evidence, not to fix):
  - Temporarily enforce a deterministic, item_id-based locking order in reserve.rb (e.g., sort line items by item_id before with_lock). If deadlocks drop, that confirms lock-order as the main issue.
  - Check for and reduce transaction scope in the checkout path if there is any outer transaction that could be contributing to locked resources.
- Longer-term considerations (once cause is confirmed):
  - Move to a single-lock operation for stock decrement if possible (e.g., a single UPDATE ... WHERE id IN (...) with a single FOR UPDATE), or serialize inventory updates to reduce cross-item contention.
  - Ensure the locking strategy is consistent across all inventory operations to avoid mixed lock modes causing cycles.

Bottom line
- The trace strongly points to a deadlock caused by concurrent reservation of multiple inventory items in a single checkout, with with_lock locking happening per item and potential lock-order inversions between two concurrent requests. Confirm this with per-item lock/order details from PostgreSQL logs, item IDs involved, and outer transaction boundaries. If confirmed, ordering locks deterministically (or reworking the reservation to a single atomic operation) would be the next steps after establishing the exact cause from the missing details.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Find why a query is slow

Given this query and its execution plan, explain where the time actually goes. Name the specific plan node responsible, say why the planner chose it, and state whether the problem is a missing index, a bad row estimate, an unnecessary sort, or the query shape itself. Propose one change and predict its effect on the plan. Do not suggest adding an index without saying which node it would eliminate. Query and plan: {{plan}}

Fill in: The query plus EXPLAIN ANALYZE output — the plan without actual timings is much weaker.

Known limits: Given EXPLAIN without ANALYZE it reasons from estimates, which is exactly what misleads when the estimates are the problem.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
SELECT articles.* FROM articles
  JOIN users ON users.id = articles.user_id
  WHERE users.organization_id = 42 AND articles.published_at > now() - interval '30 days'
  ORDER BY articles.published_at DESC LIMIT 20;

Limit (cost=98234.11..98234.16 rows=20 width=812) (actual time=3421.882..3421.889 rows=20 loops=1)
  -> Sort (cost=98234.11..98298.44 rows=25732 width=812) (actual time=3421.880..3421.884 rows=20)
       Sort Key: articles.published_at DESC
       Sort Method: top-N heapsort  Memory: 42kB
       -> Hash Join (cost=1204.55..97549.30 rows=25732 width=812) (actual time=12.441..3389.102 rows=24988)
            Hash Cond: (articles.user_id = users.id)
            -> Seq Scan on articles (cost=0.00..93122.00 rows=812004 width=812) (actual time=0.011..3102.445 rows=811233)
                 Filter: (published_at > (now() - '30 days'::interval))
                 Rows Removed by Filter: 2455102
            -> Hash (cost=1198.20..1198.20 rows=508 width=8) (actual time=12.402..12.403 rows=511)
                 -> Index Scan using idx_users_org on users (cost=0.42..1198.20 rows=508 width=8)
Planning Time: 0.284 ms
Execution Time: 3422.031 ms

Model output (unedited)

Here’s what’s happening in the plan and where the time is being spent.

What node is actually slow
- The bottleneck is the Seq Scan on articles (with the published_at filter). The scan reads about 811k rows and applies the filter published_at > now() - interval '30 days'. The time shown there (actual time ~0.011 to 3102.445 ms) dominates the total, and drives the rest of the plan (the hash join with users and the final sort/limit).
- After the scan, a Hash Join with users is performed. The inner side (users) is small (508 rows projected, 511 rows in the hash), which is cheap to hash-build, but the outer side (articles) has already incurred the bulk of cost.
- The final Sort (top-N) is used to produce the 20 rows in DESC order, but its cost is minor relative to scanning articles. The plan shows Sort input size of about 25k rows and a top-N heapsort, but the dominant time is the initial table scan.

Why the planner chose this plan
- The predicate on users.organization_id = 42 is not pushed down into the articles scan; it’s applied after joining to users. Since the articles filter is on published_at (articles.published_at > now() - interval '30 days'), and there isn’t a suitable index that can both filter by published_at and efficiently couple to the join, the planner ends up scanning articles sequentially to evaluate the predicate and then joining to the filtered set of users.
- The join type chosen is a Hash Join because:
  - The inner relation (users) is relatively small after filtering (via the index on idx_users_org), and
  - The outer relation (articles) is large; hashing the small inner side and probing with the large outer is cost-effective here.
- The plan shape is driven by the need to:
  - apply the 30-day filter on articles, and
  - then join to users to apply organization filter, and
  - finally deliver the top-20 by published_at.

Is the problem missing an index, a bad estimate, an unnecessary sort, or the query shape?
- Primary issue: missing index support on the articles filter. The planner ends up doing a full or large-range sequential scan of articles to apply published_at, rather than an index-based access that would prune far more quickly.
- It’s not primarily a bad row estimate (the estimates match the observed rows reasonably well here), and the sort is not the main culprit (the sort is small relative to the scan). The core problem is the lack of an index that can efficiently constrain articles by published_at (and potentially help with the join).

One change to try and what you can expect
- Change: add an index on articles(published_at) (btree, ideally with published_at DESC to align with the ORDER BY, e.g., CREATE INDEX idx_articles_published_at_desc ON articles(published_at DESC)).
- Predicted effect:
  - The plan would switch from a Seq Scan on articles to an Index Scan using that index for the published_at predicate.
  - This would dramatically reduce the number of articles touched before the join (instead of scanning hundreds of thousands to over a million rows, the engine would touch far fewer rows that satisfy the 30-day window).
  - Since the LIMIT 20 requires the top-most published_at values, the DESC index on published_at helps satisfy the ordering without a heavy sort, further reducing execution time.
  - The overall runtime would drop substantially (likely from several seconds to a fraction of a second to a couple seconds, depending on how many qualifying articles are in the 30-day window and how selective the organization_id filter ends up being after the join).

Notes and caveats
- If you anticipate frequent queries like this, a multi-column index that also helps with the join can help even more, e.g., an index on (published_at DESC, user_id) or a composite index (user_id, published_at DESC). However, per your instruction I’ve described one concrete change and its impact, focusing on the most direct bottleneck: an index on published_at to cut down the article scan.
- If you don’t want to rely on an index-only path (since the query selects articles.* and an index-only scan would require all article fields to be in the index), the index will still speed up access by reducing the rows read from disk and reducing the subsequent hash join cost and the sort cost.

In short
- Root cause: heavy work is the Seq Scan on articles (filtering by published_at) and the subsequent Hash Join with a large outer input.
- Plan node to blame: Seq Scan on articles (with the published_at filter).
- Best one-change suggestion: create an index on articles(published_at DESC) to enable index-based retrieval for the 30-day window and top-N ordering.
- Expected result: fewer rows read from articles, lower Hash Join cost, potential elimination of or reduction in the explicit sort, and a meaningful drop in total execution time.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Refactor a long function without changing behaviour

Refactor this function for readability while preserving observable behaviour exactly, including its error cases and return types. Produce the rewritten code, then a list of every behavioural difference you could not avoid — if there are none, say so explicitly. Do not rename anything in the public interface, do not change the signature, and do not introduce a new dependency. Flag any branch whose current behaviour looks like a bug rather than silently correcting it. Function: {{function}}

Fill in: The function, plus its tests if you have them.

Known limits: It will preserve existing bugs by design — that is the point, but it means the output is not automatically an improvement in correctness.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-mini

Input fixture
def process(order)
  if order.nil?
    return nil
  else
    if order.status == "paid" || order.status == "shipped"
      if order.items.any?
        total = 0
        order.items.each do |i|
          if i.refunded
          else
            total = total + (i.price * i.quantity)
          end
        end
        if total > 0
          return { total: total, currency: order.currency || "EUR" }
        else
          return { total: 0, currency: order.currency || "EUR" }
        end
      else
        return { total: 0, currency: "EUR" }
      end
    else
      return nil
    end
  end
end

Model output (unedited)

Rewritten code (behaviour preserved exactly):

def process(order)
  return nil if order.nil?
  return nil unless order.status == "paid" || order.status == "shipped"

  items = order.items
  return { total: 0, currency: "EUR" } unless items.any?

  total = 0
  items.each do |i|
    next if i.refunded
    total += (i.price * i.quantity)
  end

  if total > 0
    { total: total, currency: order.currency || "EUR" }
  else
    { total: 0, currency: order.currency || "EUR" }
  end
end

Behavioural differences I could not avoid
- None. The refactor preserves the original observable behaviour exactly (same return values and error cases).

Branches whose current behaviour looks like a bug (I flagged them rather than changing them)
- In the branch where order.items.any? is false, the function returns { total: 0, currency: "EUR" } ignoring order.currency entirely. In the branch where items exist but sum to 0, the function returns { total: 0, currency: order.currency || "EUR" }. This inconsistent treatment of currency between the "no items" and "items but zero total" branches looks like a likely bug.
- The original code's item-refunded check is written as "if i.refunded; else ... end" (a no-op in the if branch). I preserved that logic (i.refunded items are skipped). The empty-if style is confusing and could be mistaken for an error while reading.
- If order.items is nil (not an Array-like with any?), calling order.items.any? will raise — the original code would raise the same error. This may be surprising but I preserved that behaviour.

If you want, I can produce a variant that fixes the likely bugs (consistent currency handling; clearer refunded check; safe handling when items is nil) — but I did not change behaviour in this refactor per your instructions.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Review a change for security issues

Review this change for security defects only. Check specifically: injection through unsanitised input, authorisation checks that are missing or applied after the effect, secrets or tokens reaching logs or responses, unsafe deserialisation, and user-controlled values reaching a filesystem path, shell, or URL. For each finding give the line, the attacker-controlled input, and the concrete consequence. Do not list theoretical risks with no path from input to impact in this code. Change: {{change}}

Fill in: The change, plus any route or permission code that governs who can reach it.

Known limits: Judges only what you paste — a missing authorisation check in a parent controller is invisible to it.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
class ExportsController < ApplicationController
  skip_before_action :verify_authenticity_token, only: :download

  def download
    report = Report.find(params[:id])
    path = Rails.root.join("tmp", "exports", params[:filename])
    send_file path, filename: params[:filename]
  rescue ActiveRecord::RecordNotFound
    render json: { error: "not found", debug: params.to_unsafe_h }, status: :not_found
  end
end

Model output (unedited)

Here are security defects I can identify in the shown change, with attacker-controlled input, concrete path from input to impact, and line references:

1) Path traversal / arbitrary file disclosure (user-controlled path)
- Line: 6
  path = Rails.root.join("tmp", "exports", params[:filename])
- Attacker-controlled input: params[:filename] = "../sensitive.txt" or "../../../../etc/passwd"
- Concrete consequence: The joined path can resolve outside the intended tmp/exports directory. If the resulting file exists and is readable, send_file will serve that file to the client, enabling disclosure of arbitrary server files (e.g., sensitive configs or system files).

2) Missing authorization check prior to download (authorization without protection)
- Line: 5
  report = Report.find(params[:id])
- Attacker-controlled input: params[:id] = any valid or existing report id
- Concrete consequence: Without verifying that the current user is allowed to access this export, an attacker can retrieve exports they should not be allowed to view, leading to unauthorized data exposure.

3) Secrets/tokens leaked in error response (debug info echoes request params)
- Line: 9
  render json: { error: "not found", debug: params.to_unsafe_h }, status: :not_found
- Attacker-controlled input: any sensitive parameter supplied by the attacker, e.g., password=secret123 or auth_token=xyz
- Concrete consequence: In the not_found error path, the response includes params.to_unsafe_h, which echoes all request parameters (including secrets) back to the client. This also risks leakage if logs capture the same data. This directly exposes secrets/tokens via the API response.

4) Unsafe deserialisation
- Not observed in this snippet. No deserialisation of untrusted content (e.g., YAML/JSON) is performed here, so no concrete defect to report in this category.

5) Other notes touching the requested categories
- Line 7
  send_file path, filename: params[:filename]
  This relies on the same unsanitised filename input as in line 6. If the path can be manipulated to point outside the intended directory (as described in finding 1), the actual file retrieved may be arbitrary.

Remediation pointers (concise):
- Sanitize and validate filename to prevent path traversal (e.g., use a whitelist or map filename to a known exported file, and ensure the final path resides within Rails.root.join("tmp", "exports")).
- Enforce proper authorization before any download (e.g., verify current_user owns/has access to the Report/export, or scope the query to the user’s allowed records).
- Do not echo request parameters in error responses. Remove debug: params.to_unsafe_h or replace with safe, redacted logging.
- Consider using a secure, API-focused approach: authenticate/authorize first, then securely fetch and serve the exact allowed export.

Line-by-line summary of findings:
- Line 5 (Report.find(params[:id])): Attacker-controlled input; missing authorization guard to ensure access rights.
- Line 6 (path = Rails.root.join(..., params[:filename])) and Line 7 (send_file path, filename: params[:filename]): Attacker-controlled input; path traversal risk enabling arbitrary file disclosure.
- Line 9 (render ... debug: params.to_unsafe_h): Attacker-controlled input; risk of leaking secrets/tokens in error response.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Review a diff for correctness bugs

Act as a reviewer whose only job is finding bugs that would reach production. Review this diff for correctness only — not style, naming, or preferences. For each issue: quote the exact line, state the concrete input or state that triggers it, and state the wrong behaviour that results. If you cannot construct a failing case for an issue, do not report it. If you find no correctness bugs, say so plainly instead of inventing minor observations. Diff: {{diff}}

Fill in: The diff you want reviewed, with enough surrounding context to show how the changed code is called.

Known limits: Cannot see code outside the diff, so it misses bugs whose cause lives in a caller you did not paste.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
--- a/app/services/billing/proration.rb
+++ b/app/services/billing/proration.rb
@@ -12,9 +12,12 @@ module Billing
   class Proration
     def self.credit_for(subscription, cancelled_on:)
       days_total = subscription.period_end - subscription.period_start
-      days_unused = subscription.period_end - cancelled_on
-      return 0 if days_unused <= 0
+      days_unused = (subscription.period_end - cancelled_on).to_i
+      return 0 unless days_unused.positive?

-      (subscription.amount_cents * (days_unused.to_f / days_total)).round
+      (subscription.amount_cents * (days_unused / days_total)).round
     end
   end
 end

Model output (unedited)

Issues found: 1 correctness bug (unavoidable failing case shown). No other correctness bugs identified.

Issue 1
- Exact line(s) quoted from diff:
  - days_unused = (subscription.period_end - cancelled_on).to_i
  - return 0 unless days_unused.positive?
  - (subscription.amount_cents * (days_unused / days_total)).round

- Triggering input (concrete failing case):
  - Use date-based periods (Date objects) so differences are integer days:
    - subscription.period_start = Date.new(2023, 1, 1)
    - subscription.period_end   = Date.new(2023, 1, 31)    # 30 days total
    - cancelled_on            = Date.new(2023, 1, 16)    # 15 days unused
    - subscription.amount_cents = 10000  (i.e., $100)

- Wrong behavior that results:
  - With the diff, days_total = 30 and days_unused = (end - cancelled_on).to_i = 15.
  - The code computes days_unused / days_total using integer division: 15 / 30 -> 0.
  - credit_for returns 0 cents, even though the correct proration would credit 50% of the amount (5000 cents).
  - Expected correct result: 5000 (assuming no other rounding issues).

- Why this is a bug:
  - The original intention is to prorate based on the fraction of the period unused. The change truncates the unused duration to an integer and then uses integer division when days_total is an integer, producing zero credit in many mid-period cases. This leads to incorrect, production-persistent proration results.

- Suggested fix (conceptual):
  - Avoid integer division by ensuring floating-point division:
    - Use days_unused as a numeric that preserves fractional value (e.g., remove .to_i or convert to_f)
    - Example fix (aligned with original intent): (subscription.amount_cents * (days_unused.to_f / days_total)).round
  - If you want to preserve current guard logic, ensure the division is floating regardless of input types:
    - days_unused = subscription.period_end - cancelled_on
    - return 0 if days_unused <= 0
    - (subscription.amount_cents * (days_unused.to_f / days_total)).round

Summary
- Found a concrete correctness bug: the code now uses integer division for the prorated credit when periods are represented with integer-day differences (e.g., Date objects). This yields zero credit in many valid mid-period cancellations (e.g., 15 of 30 days unused). The failing case above demonstrates the issue.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Review a migration for production safety

Act as a DBA reviewing a migration that will run against a live table. Identify anything that takes a blocking lock, rewrites the table, or runs unbounded: adding a non-null column with a default on an old database version, adding an index without a concurrent build, changing a column type, adding a foreign key without validation deferred, and backfills without batching. For each, state the lock taken and what queries it blocks, then give the safe rewrite. Note the table size at which each stops being safe. Migration: {{migration}}

Fill in: The migration, the database engine and version, and roughly how many rows the table holds.

Known limits: Safety depends on engine version and row count — without both it answers for a generic large Postgres table.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
class AddPriorityToJobs < ActiveRecord::Migration[8.0]
  def change
    add_column :jobs, :priority, :integer, null: false, default: 0
    add_index :jobs, :priority
    add_foreign_key :jobs, :users

    Job.where(priority: 0).update_all(priority: 5)
  end
end

# PostgreSQL 14. jobs has ~40 million rows and is written to continuously.

Model output (unedited)

Below is a risk review of the migration you provided, focused on things that block, rewrite the table, or run unbounded (long-running) against a live table. For each item I list the lock type, what it blocks, a safe rewrite approach, and a practical table-size threshold indicating when it stops being safe in a live, highly-active table (like your ~40M-row jobs table).

General context for your migration
- The migration adds a new non-null column with a default, adds an index, adds a foreign key, and then backfills (in one transaction in Rails’ change method).
- On PostgreSQL, some operations lock the table aggressively (block all DML), some can be done concurrently, and backfills can be unbounded if not batched.
- With 40M rows and continuous writes, you must avoid long, single-transaction rewrites or unbatched backfills.

1) Adding a non-null column with a default on an old database version
What lock is taken and what queries it blocks
- Lock: Typically an ACCESS EXCLUSIVE lock on the table if the engine rewrites the table (older PostgreSQL behavior). In older versions, adding a NOT NULL column with a non-null DEFAULT can force a full table rewrite.
- What it blocks: All reads and writes (SELECT, INSERT, UPDATE, DELETE) on jobs while the rewrite runs. DDL on the table will also contend during the operation.
- Why this is risky here: With ~40 million rows and a live, highly active table, a full-table rewrite is effectively a maintenance window you cannot afford.

Safe rewrite (how to do it without a blocking rewrite)
- Do this in steps:
  - Step A: Add the column as NULLable with no default (no rewrite).
    - ALTER TABLE jobs ADD COLUMN priority integer;
  - Step B: Backfill in small batches, outside a single long transaction:
    - Update in chunks, e.g. in Rails: loop a bounded batch (limit e.g. 50k–100k) updating rows where priority IS NULL or priority = 0 to the desired value (or initialize to 0 first, then patch to 5 as you intended).
  - Step C: Make the column NOT NULL and set the default after the backfill completes:
    - ALTER TABLE jobs ALTER COLUMN priority SET NOT NULL;
    - ALTER TABLE jobs ALTER COLUMN priority SET DEFAULT 0;
- Why this is safer: You avoid a single long table rewrite and you reduce lock duration to batch-sized operations. Reads can continue during many of the backfills depending on MVCC and transaction boundaries.
- Table-size threshold where this stops being safe: The larger the table, the more critical it is to batch. Practically, for 40M rows, this approach is the minimum viable path. It is not safe to do a non-batched, single-step NOT NULL DEFAULT add on tables of millions of rows in a live system.

2) Adding an index without a concurrent build
What lock is taken and what queries it blocks
- Lock: Without CONCURRENTLY, CREATE INDEX acquires an ACCESS EXCLUSIVE or equivalent lock on the table for the duration of the index build.
- What it blocks: All reads and writes to jobs during the index build.
- Why this is risky here: A 40M-row table build will block for a long time, causing noticeable latency for reads and writes.

Safe rewrite (concurrent index)
- Use a CONCURRENTLY build:
  - In Rails: add_index :jobs, :priority, algorithm: :btree, concurrently: true
  - In SQL: CREATE INDEX CONCURRENTLY idx_jobs_priority ON jobs (priority);
- Why this is safer: Reads continue to work (though there may be performance degradation while the index is being built; writes can continue with some restrictions). The operation is designed to avoid full-table locks.
- Table-size threshold where this is safe: Safe for any size, but it takes longer to complete for larger tables. It is the recommended approach for 40M rows and up.

3) Changing a column type
What lock is taken and what queries it blocks
- Lock: ALTER TABLE ... ALTER COLUMN TYPE generally requires an ACCESS EXCLUSIVE lock because it is a table rewrite in PostgreSQL (or an in-place rewrite only for compatible cases in modern versions; but with risk for large tables, assume a rewrite).
- What it blocks: All reads and writes to jobs during the type change.
- Why this is risky here: On a live, heavily written table, a type change is effectively a maintenance window and can’t be done safely without a staged approach.

Safe rewrite (multi-step approach)
- Add a new column of the target type, backfill data, then drop the old column and rename the new one:
  - Step 1: Add column with the target type, nullable.
    - ALTER TABLE jobs ADD COLUMN priority_new integer;
  - Step 2: Backfill priority_new in batches from the old data (e.g., based on id ranges or time-based partitioning).
  - Step 3: Swap columns: drop the old priority, rename priority_new to priority.
  - Step 4: If needed, apply NOT NULL and DEFAULT constraints after the swap.
- Note: If you specifically need a straight ALTER COLUMN TYPE, do it only as part of a controlled maintenance window or after you’ve ensured there’s a way to perform a non-blocking swap (as above).

Table-size threshold
- Safe at all sizes only if you can perform a staged, batched rewrite (or at least a zero-downtime swap). For 40M rows, the staged approach is required; a direct in-place type change is not safe.

4) Adding a foreign key without validation deferred
What lock is taken and what queries it blocks
- Lock: If you add a foreign key with immediate validation (the default in many environments), PostgreSQL may scan existing data to validate referential integrity, which can take a long time and require a significant lock window (often ACCESS EXCLUSIVE or an equivalent strong lock during validation). If you mark the constraint as DEFERRABLE and/or NOT VALID, you can avoid the immediate heavy scan.
- What it blocks: If you validate immediately, it blocks reads/writes long enough to verify existing rows; if you use NOT VALID, you can avoid the initial heavy lock, but you still need to validate eventually.
- Why this is risky here: On a live 40M-row table with continuous writes, validating a foreign key can be expensive and blocky.

Safe rewrite (defer and validate in steps)
- Approach:
  - Step A: Add the foreign key as deferrable and not immediately validated:
    - In Rails: add_foreign_key :jobs, :users, deferrable: true, initially_deferred: true, validate: false
  - Step B: In a maintenance window, validate the constraint:
    - ALTER TABLE jobs VALIDATE CONSTRAINT jobs_users_fk
- Why this is safer: It allows you to defer the expensive validation until a low-traffic window and to minimize impact on live reads and writes. The initial NOT VALID/DEFERRED constraint addition avoids a heavy upfront full-table scan.
- Table-size threshold:
  - For very large tables (like 40M rows), this deferral pattern is highly recommended. The immediate full validation would be a long locking operation; deferring and validating in a controlled window is far safer.

5) Backfills without batching
What lock is taken and what queries it blocks
- Lock: A single UPDATE statement that touches a large portion of the table can acquire a RowExclusiveLock on each updated row; the transaction remains open until completion, creating a long-running write transaction. MVCC means reads can still proceed for other rows, but the long transaction can block competing writes and cause table bloat and locking pressure.
- What it blocks: Depending on timing and concurrent activity, subsequent writes that try to modify overlapping rows may block or be blocked; reads are generally allowed but can be slowed due to contention and bloat. The main risk is a long-running transaction that blocks autovacuum and increases table-level lock contention.
- Why this is risky here: The migration’s backfill uses update_all in one shot; with 40M rows, this is a long-running transaction that will lock portions of the table and can cause stall for concurrent writes.

Safe rewrite (batched backfill)
- Do the backfill in bounded batches with commits:
  - Example approach (pseudo-steps):
    - Set a batch_size (e.g., 50k–200k rows)
    - Loop:
      - Update a batch: Update jobs set priority = 5 where priority = 0 order by id limit batch_size
      - Commit after each batch (or run as separate small transactions)
      - Break when no more rows to update
- Why this is safer: Each batch is a small transaction, so the lock duration per batch is short, autovacuum can keep up, and you avoid a single long-running transaction that blocks other writers for hours or days.
- Table-size threshold:
  - Batched backfill is suitable for large tables (tens of millions of rows). It becomes less risky than a single huge update, and it can be resumed if the migration is paused.

Putting it all together for your specific migration
- The migration currently does in one go:
  - add_column :jobs, :priority, :integer, null: false, default: 0
  - add_index :jobs, :priority
  - add_foreign_key :jobs, :users
  - Job.where(priority: 0).update_all(priority: 5)
- What this means in PostgreSQL 14 with a live 40M-row jobs:
  - The add_column with NOT NULL DEFAULT 0 is likely to trigger a table rewrite in many environments (or at least a very long lock), blocking all reads and writes for the duration. Not safe for a live 40M-row table.
  - The add_index without CONCURRENTLY blocks all DML for the duration of the index build. On 40M rows, this is unacceptable in a live system.
  - The add_foreign_key will attempt to validate existing data; in a large, busy table this can be slow and blocking unless you defer validation.
  - The backfill with update_all is a long-running, unbatched transaction that can block writes and cause bloat; not safe for a busy system.

Recommended safe migration pattern (one example rewrite)
- Step 1: Break the migration into steps (no single big DDL with NOT NULL DEFAULT)
  - Add the column as nullable, no default:
    - ALTER TABLE jobs ADD COLUMN priority integer;
  - Create the index concurrently (so reads/writes can continue while the index is built):
    - Rails: add_index :jobs, :priority, concurrently: true
  - Add the foreign key as deferrable and not validated immediately:
    - Rails: add_foreign_key :jobs, :users, deferrable: true, initially_deferred: true, validate: false
- Step 2: Backfill in batches
  - If you plan to set existing rows from 0 to 5:
    - Batch by id, updating a fixed number of rows per iteration
    - Example (pseudo logic; adapt to your app):
      - batch_size = 100_000
      - loop do
          Job.where(priority: 0).order(:id).limit(batch_size).update_all(priority: 5)
          break if Job.where(priority: 0).limit(1).count == 0
        end
- Step 3: Make the column NOT NULL and set DEFAULT after backfill completes
  - ALTER TABLE jobs ALTER COLUMN priority SET NOT NULL;
  - ALTER TABLE jobs ALTER COLUMN priority SET DEFAULT 0;
- Step 4: Validate the foreign key constraint (in a maintenance window)
  - ALTER TABLE jobs VALIDATE CONSTRAINT ... (or Rails: run the validation step separately)
- Step 5: If you truly need a type change or a rewrite, do a staged swap
  - Add a new column with the target type
  - Backfill the new column in batches
  - Drop the old column and rename the new one
  - Add any constraints/indexes as needed

Notes on the table size threshold
- Non-null with default rewrite: not safe on any sizable table in a live system. Practically, avoid on tables larger than a few hundred thousand rows unless you can schedule a maintenance window and the hardware/IO can handle it.
- CONCURRENTLY index build: safe at any size; always preferred for live systems.
- Foreign key validation: safest to defer and validate in a controlled window; large tables should avoid immediate full validation during peak times.
- Batched backfills: safe for large tables; effectively the only sane way to backfill on a live system with heavy write traffic.

Bottom line for your migration
- The current migration as written is not safe for a 40M-row, live Jobs table in PostgreSQL 14. It will:
  - Attempt to add a non-null default column in a way that can rewrite the table (long block).
  - Create a non-concurrent index (long block).
  - Validate a foreign key on a very large table (long block).
  - Backfill in one big update (long-running transaction, high locking pressure, potential bloat).
- A safer rewrite is to split the migration into steps, use concurrent index builds, defer foreign-key validation, and batch the backfill. The safe pattern described above is what you should implement in a real environment to avoid long outages and to minimize impact on live traffic.

If you’d like, I can draft a concrete, up/down Rails migration using the batched approach (including a small script to chunk the backfill) tailored to your app’s ID distribution and how you generate IDs.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Review error handling

Review the error handling in this code. Identify: exceptions caught too broadly, errors swallowed without logging or re-raising, retries around non-idempotent operations, cleanup that will not run on the failure path, and error messages that leak internal detail to a caller who should not see it. For each, name the failure that would go unnoticed in production as a result. Do not comment on anything other than error handling. Code: {{code}}

Fill in: The code, including whatever wraps it — a swallowed error is only a problem relative to who needed to hear about it.

Known limits: Cannot know your alerting setup; 'goes unnoticed' assumes an unlogged error is invisible, which may not hold.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
def sync_customer(customer)
  attempts = 0
  begin
    attempts += 1
    response = PaymentGateway.charge(customer.token, customer.balance_cents)
    customer.update!(last_synced_at: Time.current, gateway_ref: response.id)
  rescue => e
    retry if attempts < 3
    Rails.logger.info("sync failed")
    nil
  ensure
    customer.touch(:sync_attempted_at) if customer.persisted?
  end
end

Model output (unedited)

Here are the error-handling issues I see, with the specific failure that would go unnoticed in production for each:

- Exceptions caught too broadly
  - Failure name: Silent gateway/database errors
  - What would go unnoticed: Any error raised inside the begin block (e.g., PaymentGateway::TemporaryUnavailable, ActiveRecord validation errors, or other StandardError descendants) would be rescued, retried, or swallowed without distinction. The system would not surface the specific cause.

- Errors swallowed without logging or re-raising
  - Failure name: Silent exception details
  - What would go unnoticed: The actual exception (type, message, backtrace) is not logged or surfaced. After the final retry, the code only logs "sync failed" with no details, making root cause diagnosis very hard.

- Retries around non-idempotent operations
  - Failure name: Duplicate charges
  - What would go unnoticed: If the charge fails transiently and a retry occurs, the customer could be charged multiple times for the same sync attempt.

- Cleanup that will not run on the failure path
  - Failure name: Missing/undone cleanup on failed attempts
  - What would go unnoticed: In some failure scenarios (e.g., the customer record isn’t persisted due to a prior failure), the ensure block only runs if customer.persisted? is true. This could skip the cleanup touch and leave the sync state inconsistent or untracked (sync_attempted_at not updated for a failed attempt).

- Error messages leaking internal detail to a caller who should not see it
  - Failure name: Inadequate error propagation to callers
  - What would go unnoticed: The method returns nil after failures and only logs a generic message. The caller gets no actionable error information or stack trace (and in many cases may misinterpret the result as success), while internal error details remain only in logs. A caller that should be alerted or retried at a higher level would not receive appropriate feedback.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Trace where an unexpected nil comes from

A value is unexpectedly nil or undefined at a known point. Working backwards from that point through the code given, enumerate every path that could produce it, ordered by likelihood. For each path, name the line where the value becomes nil and the condition required. Then state the single cheapest check that would distinguish between your top two candidates. Do not suggest adding a nil guard at the symptom — the point is to find the origin. Context: {{context}}

Fill in: Where you observed the nil, and the code between its origin and that point.

Known limits: Only reasons over the code you paste; if the value crosses a serialisation or job boundary, include both sides.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
# Observed: `invoice.customer_email` is nil inside InvoiceMailer#send_receipt,
# but only for invoices created through the API, never through the admin UI.

class Invoice < ApplicationRecord
  belongs_to :customer, optional: true
  before_validation :denormalize_customer_email, on: :create

  def denormalize_customer_email
    self.customer_email ||= customer&.email
  end
end

# api/v1/invoices_controller.rb
Invoice.create!(invoice_params.merge(customer_id: params[:customer_id]))

# admin/invoices_controller.rb
Invoice.create!(invoice_params.merge(customer: current_admin.impersonated_customer))

Model output (unedited)

Short answer up front
- The nil value for invoice.customer_email comes from denormalize_customer_email failing to assign a non-nil email. In this code, that happens in the line:
  self.customer_email ||= customer&.email
- There are two main ways this line yields nil, given the API create path:
  1) The API-provided customer_id refers to an existing Customer, but that Customer.email is nil.
  2) The API-provided customer_id does not refer to any Customer (or the Customer has been deleted), so customer is nil (or missing), and customer&.email is nil.

Paths, ordered by likelihood (given “only for API-created invoices” and the code shown)

Path #1: API path with a valid Customer that has no email
- How the nil arises: There is a Customer for the given id, but that Customer.email is nil. The denormalization runs during create, and self.customer_email ||= customer&.email evaluates to nil because the right-hand side is nil.
- Where the value becomes nil: 
  - The line: self.customer_email ||= customer&.email
- Condition required for this path:
  - invoice.customer_id corresponds to an existing Customer, AND that Customer.email is nil.
- Example scenario in code terms: Customer.exists?(id: invoice.customer_id) is true and Customer.find(invoice.customer_id).email is nil (or blank).

Path #2: API path with a non-existent (or deleted) Customer
- How the nil arises: There is no Customer for the given id, so customer is nil; customer&.email is nil; the ||= keeps customer_email as nil.
- Where the value becomes nil:
  - The line: self.customer_email ||= customer&.email
- Condition required for this path:
  - invoice.customer_id is nil or does not correspond to any existing Customer (Customer.find_by(id: invoice.customer_id) returns nil, or the Customer has been deleted after the invoice was created).
- Example scenario in code terms: Customer.exists?(id: invoice.customer_id) is false.

Notes on other possibilities (less likely given the observed path)
- Admin flow: If the impersonated_customer had nil email, the same denormalization could yield nil for an admin-created invoice too, but the symptom says it’s only seen for API-created invoices. It’s worth noting but not ranked among top two.
- Any post-denormalization override: Not present in the shown code; the line above is the only place in this snippet that sets customer_email during create.

Cheapest check to distinguish top two candidates
- Do a single existence check on the associated Customer:
  - Check: does a Customer exist with the invoice’s customer_id?
  - Ruby/Rails form: Customer.exists?(id: invoice.customer_id)
- How to interpret the result:
  - If exists? returns false, you’re in Path #2 (no such customer; customer is nil).
  - If exists? returns true, you’re in Path #1 (customer exists but may have nil email; you’d then check Customer.find(invoice.customer_id).email to confirm).

What to do next (quick verification steps)
- Inspect the invoice that ended up with nil email:
  - Retrieve: inv = Invoice.find(id)
  - Check: inv.customer_id
  - Check existence: Customer.exists?(id: inv.customer_id)
  - If exists? is true, load the customer: c = Customer.find(inv.customer_id); examine c.email
  - If exists? is false, there was no associated Customer at the time of denormalization (or it was deleted)

This targets the origin efficiently: the single existence check tells you whether you’re on the “non-existent customer” path vs the “existing customer but no email” path, which are the two plausible routes given the code.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Turn a vague bug report into a reproduction

Act as an engineer triaging a report too vague to act on. Produce: the minimal sequence of steps that would reproduce it if your leading hypothesis is right, the specific environment details that matter and why, the three questions to ask the reporter ranked by how much each narrows the search, and what you would check in logs or telemetry before asking a human anything. Do not guess at a fix. Report: {{report}}

Fill in: The report exactly as the user wrote it, without your own interpretation layered on.

Known limits: Its hypothesis is only as good as the surrounding context you give it; with no system detail it asks broader questions.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-nano

Input fixture
"The editor keeps losing my work. I wrote a long post yesterday and when I came
back today it was mostly gone, only the first paragraph was there. This has happened
twice now. I'm on a Mac. Please fix, I'm not going to keep writing here if this keeps
happening."

Known: the editor autosaves a draft every 20 seconds via PATCH /articles/:id.
Drafts are also stored in localStorage as a fallback. No related errors in Sentry.

Model output (unedited)

Here’s a focused triage plan you can use to investigate the vague report, assuming the leading hypothesis is a race/concurrency issue between autosave patches and the local draft fallback.

Leading hypothesis (for reproduction steps)
- A race between autosave PATCH requests (every 20s) and the localStorage fallback draft can cause the editor to overwrite a newer in-editor state with an older draft stored in localStorage, especially when there are cross-tab or cross-tabbed edits on the same article. The symptoms described (long post yesterday, today only the first paragraph remains) fit a scenario where an older draft state is re-applied after a newer change has been made.

Minimal reproduction steps if this hypothesis is correct
- Prereqs: Use a Mac browser (Chrome or Safari) with a single article id in edit mode. Ensure you have a long post (multi-paragraph) to clearly see content loss.
- Steps:
  1) Open the article in editor for article_id = X.
  2) Type a long post with multiple paragraphs (e.g., 800–1500+ words).
  3) Wait for at least one autosave cycle ( ≥ 20 seconds ). Let the PATCH /articles/:id request occur and complete.
  4) In a second tab (same browser or another tab in the same browser/profile) open the same article (same article_id X) and make a small edit or simply keep the tab idle for an autosave cycle to happen there as well.
  5) Return to the first tab and reload the editor (or briefly switch away and come back). Observe whether the content reverts to a much earlier draft (e.g., only the first paragraph remains) or otherwise differs from what you see in the second tab.
- Expected outcome if the leading hypothesis is right: after a second autosave in another tab and subsequent reloading, the first tab shows content that looks like an older draft (potentially truncated to the first paragraph) rather than the latest in-editor state.

Environment details that matter and why
- Browser and version (e.g., Safari on macOS, Chrome on macOS, or Firefox on macOS) and macOS version.
  - Why: Different browsers have different behavior around contenteditable, localStorage, and tab/background activity. Race conditions may manifest differently (or only in certain browsers).
- Article editing workflow specifics
  - Article ID being edited, and whether the same article is edited in multiple tabs, windows, or devices simultaneously.
  - Why: Cross-tab edits are a common source of write conflicts with autosave loops and localStorage fallbacks.
- Network status during edits
  - Online vs offline, intermittent connectivity, and any proxy/VPN behavior.
  - Why: LocalStorage fallback is used when the server patch can’t be saved; flaky network can increase chances of stale data being re-applied.
- Editor/content format
  - Rich text vs. markdown vs. plain text; size/length of content; presence of embedded content (images, embeds, formatting).
  - Why: Patch payload shape and how content is serialized may affect truncation or partial patches.
- Extensions or privacy tools
  - Any extensions that block storage, modify network requests, or alter tab activity.
  - Why: Extensions can interfere with localStorage, patch requests, or tab lifecycle.

Three questions to ask the reporter (ranked by narrowing power)
1) Are you editing the same article in more than one tab or device at the same time?
   - Highest impact for narrowing to cross-tab race conditions.
2) Do you notice any PATCH /articles/:id requests failing, timing out, or being delayed around the times this happens? Were you offline at any point?
   - Helps confirm network-related queuing or stale-replay issues that could overwrite the draft.
3) Which browser(s) and versions are you using, and does the issue reproduce across different browsers?
   - Narrows down environment-specific issues (Safari vs Chrome, etc.) and helps reproduction across stacks.

What you would check in logs or telemetry before asking a human
- Frontend (client) activity
  - Patch request stream for this article_id: timestamps, request payloads, response status, and response content.
  - Compare the patch payload content with what is shown in the editor at the time of patch: is the patch sending the full current content, only a delta, or is it truncating content?
  - LocalStorage state for drafts: draft_article_<id> keys, their content, and last_saved_at timestamps. See if a stale localStorage draft is being reapplied after a patch.
  - Page load flow: whether the UI ends up rendering server content, or content from localStorage, or a mix after reload.
- Server / backend
  - Patch logs for PATCH /articles/:id: user_id, timestamp, payload delta, and the resulting stored content version.
  - Article content version history (audit log): any out-of-order patches, reverts, or concurrent edits that could cause content rollback.
  - Any app-level merge/conflict resolution logs if a patch overlaps with another patch or an edit state.
- Telemetry correlation
  - Link patch requests to a single user session (correlation IDs) to see if the same user’s patches are being overwritten by an earlier draft in another tab.
- If available, browser/console signals
  - Pre-existing console warnings or errors related to localStorage quota, contenteditable events, or failed fetch/XHR for PATCH.
- Why this helps before a human asks for more info
  - You can determine if the problem is due to cross-tab concurrency, network instability, or a mismatch between the UI state and stored drafts. It also helps decide whether to attempt a targeted device/browser test, or to reproduce in a controlled environment (same article, multiple tabs) before escalating.

Notes
- Do not propose a fix here. The goal is to confirm whether the race between autosave and localStorage/existing open tabs is happening, and to gather concrete signals that point to a root cause.
- If the above steps reproduce the issue, gather the requested telemetry and share with the team (frontend, backend engineers) to determine the exact write path and ordering logic that could lead to a regression.

If you want, I can tailor the reproduction steps to a specific stack (e.g., if you know the exact API payload shape or the frontend framework), or help you draft a quick telemetry query or log parsing checklist to run in your environment.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)
Curated

Write a failing test that reproduces a bug

Given this bug report and the relevant code, write the smallest test that pins the CORRECT behaviour — so that it fails against the code as it stands today and passes once the bug is fixed. Use the test framework already present in the code. State in one line what the test asserts, and state explicitly which line of the current code makes it fail today. If the report lacks the detail needed to reproduce it, list what you would need to ask instead of writing a test. Report and code: {{report}}

Fill in: The bug report as written, plus the code path you believe is involved.

Known limits: A vague report yields a test for the wrong thing — the prompt asks it to demand detail instead, so expect questions back sometimes.

Run evidence

Sandbox run 2026-07-30 · openai/gpt-5-mini

Input fixture
BUG: Users in Sydney report their weekly digest arrives on Sunday, not Monday.
Users in London get it on Monday correctly. Started after we shipped timezone support.

def self.due_now?(user, now: Time.current)
  local = now.in_time_zone(user.timezone)
  local.monday? && local.hour == 9
end

# Called from a cron that runs hourly, on the hour, in UTC.
# Tests use RSpec. Sydney is UTC+10 or UTC+11 depending on DST.

Model output (unedited)

I can’t write a reliable failing test from the report alone — there are too many unspecified runtime/timezone details that affect the outcome. Before writing a deterministic test I need the following facts:

1. The exact string/value stored in user.timezone for Sydney users (e.g. "Australia/Sydney" vs a fixed offset like "+11:00").  
2. The exact UTC hour(s) the cron runs when the digest job is checked (e.g. runs at 22:00 UTC and 23:00 UTC or only at 23:00 UTC). The report says “runs hourly, on the hour in UTC” but we need which of those runs produced the unintended Sunday send.  
3. A specific example timestamp (UTC) when a Sydney user actually received the digest (and the corresponding local time reported by the user). A single concrete example (UTC and local) will let me reproduce.  
4. Whether Rails’ Time.zone / application time zone is set anywhere during cron execution (what is Time.zone while the cron runs?). In Rails, Time.current == Time.zone.now so this matters.  
5. The Rails version (behaviour of in_time_zone / TimeWithZone has changed across versions) and whether tzinfo-data is present in the environment.  
6. Whether the production server had recently switched DST for Australia/Sydney (i.e. whether the incident occurred near DST boundaries).

What I would assert once I have the above:
- One-line behaviour to pin: “A user with timezone 'Australia/Sydney' should have due_now?(user, now: <UTC time when cron runs>) return true exactly when the user’s local time is Monday 09:00.”

Which current line in the code is likely responsible for the bug:
- The conditional line local.monday? && local.hour == 9 is the one that implements the check; if the inputs to in_time_zone() or the meaning of Time.current are wrong the condition will evaluate at the wrong UTC hour and cause the send to happen on local Sunday. Concretely, the line that can cause the observed failure is:
  local = now.in_time_zone(user.timezone)
  local.monday? && local.hour == 9
(the second line — local.monday? && local.hour == 9 — is where the decision is made and thus where the bug manifests).

Suggested minimal RSpec test I will write as soon as you supply the missing details (example shown with assumptions you must confirm):

- Preconditions I need you to confirm: user.timezone == "Australia/Sydney", cron run that mis-sent happened at 2024-12-01 22:00:00 UTC (example), Rails app Time.zone is UTC during cron.

Hypothetical test (replace the example UTC time with the incident time you provide):

it "returns true when the UTC cron run corresponds to Monday 09:00 in Australia/Sydney" do
  user = User.new(timezone: "Australia/Sydney")
  # Example problematic UTC time that should correspond to Monday 09:00 in Sydney
  now = Time.utc(2024, 12, 2, 22, 0, 0)  # 2024-12-02 22:00:00 UTC == 2024-12-03 09:00:00 AEDT
  expect(MyClass.due_now?(user, now: now)).to be true
end

One-line assertion: "At the given UTC instant the method should return true because that UTC instant equals Monday 09:00 in Australia/Sydney."

When you provide the concrete timezone string and the offending UTC timestamp (and confirm app Time.zone during cron), I will produce the smallest failing test that runs against the current code and will fail today and pass after fixing the underlying timezone handling bug.

Editorial review (independent AI agents)

  • Accuracy — pass (editorial-accuracy/google/gemini-2.5-flash)
  • Safety — pass (editorial-safety/anthropic/claude-haiku-4.5)
  • Usefulness — pass (editorial-usefulness/openai/gpt-5-nano)

Get the daily AI digest

One email each morning — what shipped in AI, and the prompts we tested. No filler.