improvemyprompts.
← Back to the checker

Learn from results, not prettier prompts

The research question is simple: which small repairs help people get the result they intended, and when should we leave their request alone?

What a check records

API and MCP checks can record a pseudonymous case identifier, algorithm version, task category, repair combination, disposition, word counts and evaluation criterion. They do not store the prompt or answer. You can opt out with research: false. The homepage’s browser checker does not send cases to this dataset.

What counts as an outcome

A copy, a click or an accepted rewrite is not evidence of better output. Clients can report a paired comparison using the same model, inputs and settings. Choose the criterion before testing. Tell us whether each output passed, failed, was mixed, or was not tested. Optional token and latency counts must be actual measurements. Claims about the setup and results are client reports; this service has not independently observed the underlying model run.

How weekly drafts are made

Weekly snapshots separate task, algorithm version, repair combination, criterion, model and evidence source. Comparable pairs require pass/fail outcomes, declared same conditions and a matching criterion selected before the check. A draft shows a cohort only after 30 comparable reports. That is an editorial minimum, not statistical significance. If there is too little evidence, the draft says so.

Feedback is voluntary, users are not authenticated research participants, and the service cannot prove independent identity. Reports can be biased or manipulated. Repair combinations prevent causal attribution to a single rule. These are observational field notes, not randomized experiments or performance guarantees.

Observe gaps before proposing changes

Private snapshots include policy/action/assessment breakdowns and missing-report counts. No feedback means unknown, not success, failure or confirmed abandonment. Corrected reports retain history; snapshots retain their original cutoff and do not retrospectively change. Current synthetic development tests are not a protected holdout and do not demonstrate downstream task improvement.

How we improve the algorithm

Look for repeated regressions, inspect explicitly donated examples in private, propose a focused candidate change, then compare it with the existing version and native host on independently controlled task episodes. Preserve failures as well as successes. A pattern is a hypothesis for testing, not an automatic algorithm update. Candidate proposals are versioned, evaluated and gated by separate owner and evaluator signatures; bounded rollout, rollback and emergency disable preserve control.

Donated examples

Donation is optional and separate from outcome reporting. A client must obtain explicit human consent for each example, and the user must confirm redaction. Automated checks catch some common sensitive patterns but cannot guarantee anonymity. Donations are quarantined for private review and excluded from generated articles. Donating does not authorize public quotation or model training.

Weekly jobs produce private drafts with evidence snapshots. They do not automatically publish articles.