How hotels can measure recommendations in ChatGPT
A hotel can begin measuring ChatGPT recommendations by asking a fixed set of realistic traveller questions in fresh, non-personalised chats, retaining every answer and recording whether the property is proposed for the stated trip. Repeat the questions under recorded conditions and calculate results against the valid answers collected. This produces a benchmark for those questions, not a measure of all ChatGPT users or hotel bookings.
Published by AI See You. Last reviewed 10 September 2026.
The manual pilot below is intended for a hotel marketing team with a spreadsheet and access to ChatGPT. It is an editorial method published by AI See You, not an OpenAI standard, a hands-on product test or a description of AI See You’s production scoring. Its value is a record another person can inspect and repeat.
Define the trip before asking about the hotel
Start with a decision the hotel cares about: whether it appears among the options for couples visiting without a car, for example. Choose traveller needs the property can credibly serve, including important constraints. A recommendation for a trip the hotel cannot accommodate is not a useful win.
Keep discovery questions separate from questions that name the property. “Where should we stay?” tests whether ChatGPT proposes the hotel without being given its name. “Is Harbour House a good hotel?” tests a prompted assessment. Combining their results would flatter a discovery measure because some questions have already introduced the brand.
The examples below form an illustrative Sydney panel. Replace the destination and constraints with ones relevant to your hotel, then save the exact wording before collection. These are invented prompts, not evidence of traveller search volumes.
Scenario | Example discovery question |
|---|---|
Weekend without a car | “We are two adults visiting Sydney for a weekend without a car. We want a quiet hotel with easy public transport to the harbour. Which hotels would you suggest, and what are the trade-offs?” |
Family stay | “We are visiting Sydney with two children aged six and nine. We need accommodation where the four of us can sleep in one room and reach major attractions without driving. Which hotels should we consider?” |
Conference visit | “I am attending a conference at ICC Sydney and want a hotel within a 15-minute walk, with a desk suitable for working. Which hotels would you suggest?” |
These questions examine general suitability. A separate dated-trip panel could include exact dates, party size and budget with currency. Decide that upfront: availability-sensitive questions and general recommendations are different tests. If you change dates later, label that change rather than treating the new trip as identical.
Set up a baseline you can describe
Use ordinary ChatGPT rather than a custom GPT, an existing planning conversation or a project containing hotel files. Open a new Temporary Chat and choose non-personalised responses before starting. OpenAI’s Temporary Chat documentation says this mode does not use memory, custom instructions or plugins. Personalised Temporary Chat is a different option; the word “temporary” alone is not enough to describe the test.
For this pilot, select Search for every question. OpenAI documents both automatic and manually selected search. Record the selected mode and any visible search activity. If Search is unavailable, log the problem and retry later instead of silently substituting another mode.
Keep the collection device, network and location permissions consistent. Record your actual collection country, the visible model or mode label, account/workspace and date with time zone. ChatGPT can use approximate IP location and optional device location. Writing “I am travelling from London” provides a scenario; it does not prove you have reproduced a user physically searching from London.
Non-personalised chat does not make every response identical or remove all context. OpenAI notes limited safety-related exceptions. If you cannot establish the non-personalised setting, keep that collection separate from this baseline. Record other unexposed settings as unknown. Do not infer a hidden model version from the answer’s style.
Collect a small pilot without selecting the winners
One manageable starting exercise is the three questions above, asked once each on three collection dates: nine scheduled observations. That is a practical rehearsal of the process, not a statistically sufficient sample or a prescription for every hotel. Choose the dates in advance and use a fresh non-personalised chat for every observation.
Copy the saved question exactly. Retain the first completed answer, including qualifications and hotels you did not expect. Do not regenerate until your property appears. If a capture fails, preserve the failed attempt and link any retry to it; the retry should not become an unexplained extra observation.
Keep follow-up questions out of the baseline. Asking “What about our hotel?” may be useful for diagnosis, but it changes the conversation. Save that exchange as a separate investigation rather than adding its answer to the discovery results.
Temporary chats do not normally remain in history. Copy the complete answer into your evidence file before closing it, together with cited URLs and a screenshot where useful. Store one spreadsheet row per scheduled observation and link it to the capture. A score without its answer is difficult to audit.
Record field | What to save |
|---|---|
Observation identity | Scenario ID, exact prompt, collection round, timestamp and time zone |
Collection conditions | ChatGPT interface, visible model/mode, Search selection, non-personalised state, actual country/network and location permission |
Capture status | Completed answer, clarification request, error or unavailable capture; retry reference if applicable |
Hotel outcome | Correct property identity, classification, relevant quotation and any conditions |
Competing options | Other properties proposed and the needs or trade-offs attached to them |
Evidence | Full answer, actual inline citation URLs, separate Sources-panel links and reviewer notes |
Classify the answer before calculating a percentage
Write the rules before reviewing the results. For this pilot, count a recommendation when the answer proposes the correct property as an option for the stated trip. Count each property at most once per answer, even if its name appears several times. A link to the hotel website alone is not a recommendation.
Use a primary outcome of recommended, mentioned only, discouraged, absent or unresolved. “Unresolved” is useful when identity or suitability cannot be judged confidently. A group name without a property location should not automatically receive credit for your hotel.
Conditions need judgement. “Consider Harbour House, but check room availability” can remain a recommendation with an availability caveat. “Harbour House is attractive, but it cannot fit four people in one room” should not count as a recommendation for the family scenario above. Preserve the wording rather than allowing a positive adjective to override the constraint.
Have a second person review unclear classifications. If you still cannot resolve an answer, report that uncertainty. Do not quietly treat it as a loss or remove it because it makes the headline less convenient. The distinction between visibility and recommendation intelligence explains why these outcomes need separate treatment.
Calculate within each scenario first
For a simple pilot measure, divide the valid answers recommending the property by all valid answers to that scenario. A valid answer here is a completed accommodation answer responsive to the question, including one that recommends other hotels. A technical failure or a request for clarification does not qualify under this rule and must be reported separately. Keep unresolved hotel classifications visible rather than forcing them into a definitive numerator.
Suppose, hypothetically, the weekend scenario produces three valid answers. Your hotel is recommended in one, mentioned only in another and absent in the third. Report “recommended in one of three answers for this weekend scenario”, optionally alongside approximately 33%. Do not report “33% of travellers choose us”. One additional recommendation in this tiny sample would change the result to two of three, which illustrates how unstable a small percentage can be.
If the family scenario has only two valid answers because one capture failed, report that denominator and the missing observation. Do not call the failed capture an omission of the hotel. For an unresolved classification, show a range if useful: one confirmed recommendation, one confirmed non-recommendation and one unresolved result among three valid answers means the recommendation rate could be between one and two of three, pending review.
Keep the three scenarios visible rather than rushing to a single hotel score. Equal numbers of observations do not mean that the traveller needs are equally common or commercially valuable. More repeats can reveal variation within your chosen questions, but they do not establish how often real people ask them.
Inspect sources and act on checkable problems
Open the actual citations attached to statements about the hotel. OpenAI says the Sources panel can also contain other relevant links, so retain that distinction. A source listing an old property name or an inaccurate room configuration is a specific problem to investigate. A citation does not prove why ChatGPT selected a hotel.
Create an action record with the disputed claim, source URL, correct evidence, owner and date checked. Correct your own material where necessary and seek corrections from third parties when appropriate. Keep hypotheses separate: a recurring association with nightlife might merit investigation, but it does not prove that changing one paragraph will win quiet-weekend recommendations.
Run the unchanged panel again after a planned interval, retaining the earlier evidence. Mark changes in the interface, model label, questions or conditions. A rise after a website edit is a signal to examine, not proof that the edit caused the rise. Referral sessions, enquiries and bookings require separate evidence.
A spreadsheet is enough while the team can retain and review every answer. If maintaining that record becomes difficult, the hotel tools guide covers alternatives. AI See You’s methodology describes its separate approach to recommendation measurement across measured travel markets.