How hotel groups should measure AI recommendation performance

A hotel group should measure recommendation performance through a stable set of traveller scenarios, resolved property identities and clearly defined outcomes. Report results by property, destination, traveller need and AI surface before combining them. Keep expansion of the measurement programme separate from movement in the comparable sample, and keep recommendation evidence separate from enquiries and bookings.

Published by AI See You. Last reviewed 10 September 2026.

This is a proposed management framework for hotel groups. It is not an industry-standard scoring formula or a specification of every AI See You feature. AI See You publishes the guide; its own production approach is described separately in the measurement methodology.

Decide what the monthly meeting needs to answer

A group-wide number is rarely enough to assign work. If headquarters learns that AI presence increased, it still needs to know which properties gained, for which travellers and whether the change reflects the captured answers or a change in collection.

A useful monthly meeting should be able to identify a property or traveller segment worth investigating, inspect the supporting answers and agree who will check the facts. Regional teams need a view relevant to their markets. The board needs an honest account of the strength and limits of the evidence.

Write those decisions into the measurement brief. A programme designed to assess recommendation competition should not quietly become a count of corporate brand mentions because those are easier to collect.

Establish the measurement unit

Define an observation as a captured answer to a recorded question under specified collection conditions. At minimum, retain the destination, traveller need, AI surface, collection time, origin-market context, language and the exact question. Record available model or mode information and any limitations of the collection method.

Keep the AI surfaces separate. ChatGPT, Claude, Gemini, Perplexity, Google AI Overviews, Google AI Mode and Microsoft Copilot are distinct customer-facing experiences. An API response is not automatically equivalent to the consumer interface. Ask the measurement provider how each result is obtained before presenting the coverage as like-for-like.

Google documents that AI Overviews do not trigger for every search. No Overview is a different event from an Overview that appeared and omitted the hotel. A failed capture is different again. Report feature appearance among successfully checked searches separately from recommendations among valid generated answers. Otherwise, a fall in how often the feature appears can be mistaken for a change in its hotel selections.

Resolve the portfolio before calculating its performance

Maintain a property register with current names, former names, brand membership, destination and relevant website addresses. Give each property a stable internal identifier. The point is continuity: a rebrand should not automatically become a lost old hotel and a newly discovered replacement.

Document how group-only mentions are handled. If an answer names the hotel brand without choosing a location, record that at brand level. Do not award every property a recommendation. Likewise, a citation to a shared booking domain needs interpretation before it is attributed to one hotel.

Use difficult examples during setup and review exceptions periodically. Similar names, mixed-use developments and hotels that change affiliation can defeat a tidy spreadsheet. Someone should own those corrections and retain a record of decisions.

Keep a stable core and a separate expansion sample

The core sample is the set used for movement reporting. It should preserve comparable destinations, traveller needs, question definitions, surfaces and collection conditions. The expansion sample explores new properties, markets or questions without silently changing the baseline.

Suppose September includes additional destinations where the group has strong properties. An improved overall percentage might reflect that addition rather than stronger recommendation performance in the original markets. Show the original sample’s movement separately from the new coverage.

This does not mean the programme must remain frozen forever. Revise it when commercial priorities change, but record the version and establish a new baseline where comparability is broken. Keep a short change log accessible from the report.

AI See You’s methodology similarly requires comparable measurement panels for its monthly Pulse movement. Buyers should ask any provider to explain which observations are eligible for a period-to-period comparison and which have been excluded.

Define recommendation credit and expose the denominator

Agree how to classify a recommendation, a contextual mention, discouragement and a conditional answer before scoring the sample. The visibility and recommendation explainer shows why these outcomes should not be interchangeable.

One practical measure is the number of valid sampled answers recommending a property divided by the number of valid sampled answers in the relevant scenario set. Count a property at most once per answer under this illustrative rule. Define a valid answer in advance, including how clarification requests and other non-accommodation responses are handled. Show both counts alongside the percentage. Retain unresolved property or recommendation classifications for review rather than forcing them into a definitive rate.

Record unavailable or failed observations separately. Dropping them from a valid-response denominator does not make them disappear: the report should also show the expected sample, completed observations and missing results. A market with poor capture coverage should not look equally well evidenced as one with complete coverage.

Choose portfolio aggregation deliberately

Consider two hypothetical hotels measured against their own relevant scenario sets.

Property

Answers recommending it

Valid sampled answers

Recommendation rate

Hotel North

8

10

80%

Hotel South

20

100

20%

The average of the two property rates is 50%. Pooling all observations produces 28 divided by 110, approximately 25.5%. Both calculations are arithmetically correct. The first gives each hotel equal weight; the second gives greater weight to the hotel with more observations.

Neither should be called the group’s universal recommendation performance without qualification. The scenarios may differ, and the small sample for Hotel North is especially uncertain. The example explains weighting, not a reliable estimate for a real portfolio.

If the group uses a commercial weighting, such as an agreed property priority, set and disclose it before reviewing the outcome. Keep the unweighted property results visible. Otherwise, a change in weighting can conceal a deterioration in an important market.

Build the report from detail upwards

Start with the property and traveller segment. Then show the destination context and the separate surface results. Only after those views should the report present a portfolio summary.

A useful report entry contains the scenario, the relevant comparison period, the size and completeness of the sample, captured examples and the next investigation. “Hotel South lost family recommendations on one surface” is more actionable than “AI score declined”, but it still needs evidence and an assessment of ordinary answer variation.

Avoid ranking properties with very different samples as if the comparison were controlled. A city conference hotel and a rural leisure property may have different traveller questions, competitors and seasonality. Their local context can be more useful than their place in a national table.

Separate response evidence from commercial outcomes

Recommendation measurement asks what the sampled AI answers proposed. Website analytics asks what visitors did on the hotel’s own properties. Enquiries, booking records and customer research add further evidence. They should inform one another without being collapsed into a single invented attribution figure.

Google’s generative AI performance report can show link impressions from supported Google Search AI features. It does not establish how many travellers saw a recommendation in every AI assistant, and its aggregation rules matter when comparing property-level pages with domain totals.

Where referral data is available, report it alongside the recommendation findings with its own definition. Missing referral information does not prove that no traveller used AI; an observed referral does not prove the measurement programme created an incremental booking.

Assign corrections and review what changed

Use each material finding to create a bounded investigation. The property team verifies facilities and access details. Marketing checks the relevant website copy and cited sources. Technical teams investigate an established access issue. A measurement owner confirms that the apparent movement survives a comparability and evidence review.

Record the action and the date, then observe the unchanged core sample again. If performance improves, describe the improvement without automatically attributing it to the last content edit. Several changes can occur at once, including the AI platform itself.

Keep unresolved findings on the agenda until the evidence or limitation is clear; do not present an investigation as a completed improvement. The hotel-group tool selection guide turns these requirements into questions for a vendor demonstration.