Testing record · June 30, 2026

We kept the receipts.

We reviewed 100 prompt variants for structure, selected 18 finalists, and published six winners. The complete output test below covers three representative tasks. This is an internal editorial benchmark, not an independent scientific study.

Method

ModelOpenAI GPT-5 Codex session
RunsOne per prompt
Cases3 A/B pairs
ScoringBinary, 5 checks each

Each pair received the same fictional facts. The generic version used a realistic one-line request. The structured version used the corresponding Field Guide template. A check passed only when the visible output satisfied it without charitable interpretation.

Limitations: one model, one run, fictional data, and an internal evaluator. Results show instruction-following in these cases only. They do not establish universal superiority or replace human compliance review.

Results

Variants reviewed100
Prompts published6
Generic pass rate40%
Field Guide pass rate93%

Case 01 · Instagram caption

Verified property facts

Two bedrooms, two bathrooms, 1,180 square feet, private balcony, one parking space, open Sunday from 1–3 PM. No renovation date, neighborhood ranking, or scarcity claim supplied.

Generic · 1/5

Prompt: Write an exciting Instagram caption for my new listing.

Step into luxury in this beautifully renovated gem in one of the city’s most sought-after neighborhoods. This one won’t last—schedule your private tour today!

Passed: clear CTA. Failed: factual discipline, requested format, supplied context, prohibited-language control.

Field Guide · 5/5

Prompt: Instagram caption template with facts, length, tone, blacklist, and audit rules.

Two bedrooms, two baths, and 1,180 square feet—with a private balcony and one parking space included. Join us Sunday from 1–3 PM for a closer look at the layout and details. Send a message for the full listing information.

Passed all five checks.

Case 02 · Lead follow-up SMS

Known lead context

Jordan downloaded the 14 Oak Street brochure yesterday. No prior conversation, timing, financing, or budget information is known.

Generic · 2/5

Prompt: Write a follow-up text to this real-estate lead.

Hi Jordan! Just checking in to see if you’re still interested in 14 Oak Street. I’d love to schedule a quick call and help you find your dream home. What time works for you?

Passed: uses the property, asks a question. Failed: pretends ongoing contact, generic language, high-effort meeting request.

Field Guide · 5/5

Prompt: Lead SMS template with known context, relationship, length, and prohibited phrases.

Hi Jordan, I saw that you downloaded the brochure for 14 Oak Street yesterday. Was there a specific feature or detail you wanted to look at more closely? I’m happy to send the most useful information first.

Passed all five checks.

Case 03 · Seller update

Raw notes

Four showings this week. Two buyers mentioned the second bedroom felt small. One buyer asked about parking. No offers. Online saves increased from 18 to 24. Recommendation: keep price unchanged through the weekend and reassess Monday.

Generic · 3/5

Prompt: Turn these notes into a positive seller update.

Great news—we had strong activity this week with four showings and growing online interest. Buyers are engaging with the listing, and the increase in saves shows momentum. While we do not have an offer yet, I’m confident we are moving in the right direction. Let’s stay the course!

Passed: includes activity, stays concise, gives direction. Failed: overstates saves and omits material feedback.

Field Guide · 4/5

Prompt: Seller-update template requiring facts, interpretation limits, recommendation, and one decision question.

Subject: Weekly listing update

We had four showings this week. Two buyers said the second bedroom felt small, one asked about parking, and online saves increased from 18 to 24. We do not have an offer yet.

The repeated bedroom feedback is worth watching, but four showings are not enough to establish a market-wide conclusion. My recommendation is to keep the price unchanged through the weekend and reassess Monday after we have the next activity. Are you comfortable holding the current price until that review?

Failed one check: “repeated feedback is worth watching” is reasonable interpretation, but the rubric required avoiding any pattern language from only two comments.

Why the templates are structured this way

OpenAI prompt engineering guidance Google Gemini prompting strategies Anthropic prompting best practices