Vamsi Alla

Work

Watch the 2-minute demo

Recorded on the live app with real answers: a cited answer in seconds at two hotels, a hotel rule overriding a loyalty perk, a manager escalation, a refusal with suggestions, and the guidance a first-time visitor sees.

Try the live app Narration: AI voice (Piper, open source) · short captions in the video, full narration captions with the CC button · no sign-in needed

Hotel Operations Knowledge Assistant

Correct, property-specific policy answers for front-desk staff in seconds, with the source for every fact.

This is a simulated client engagement: a fictional hotel group with a beach resort and a mountain lodge. The documents are synthetic, written to mirror real hotel operations. Quality, speed and cost are measured; the business value figures are projections from stated assumptions.

On the live app, a front-desk agent clicks Can a guest bring their dog? at Beach Resort and a cited answer streams in with its source chip and related questions
The live app: one click on a suggested question, a cited answer from the hotel's own policy.

The problem

A guest at the front desk asks, "Can I bring my dog?" At the beach resort, the answer is yes, up to 40 lbs, $75 per stay. At the mountain lodge, it's yes, any size, $50 per night, capped at $200. And an outdated 2024 document on the shared drive still says no pets at all.

Staff have to find the right document, for the right hotel, while the guest waits. That happens dozens of times a day: pets, late checkout, parking, fees, the spa, dining, events. Wrong answers end in refunds, complaints, and calls to the manager.

What I built

  • Property-aware chat: staff pick their hotel; answers use that hotel's rules plus brand-wide rules, never another hotel's, and each hotel keeps its own conversation.
  • A source for every fact, streamed in about a second, with chips showing the document, section and the policy's effective date.
  • Refusal instead of guessing, and off-topic questions refused before any API call.
  • Corrections, not agreement: if staff state a wrong fee or time in the question, the answer corrects it from the documents.
  • "Check with a manager" notes for refunds over the front-desk limit, safety and repeated complaints, quoting the rule word for word.
  • Pages for managers: Insights (what staff ask, helpful rate, unanswered questions as policy gaps) and a Document library showing what was indexed, excluded or needs OCR.
  • Public demo protections: per-session and daily limits, friendly errors, output escaping, a private usage log and an owner-only admin page.
Gold member late checkout answer at Beach Resort: capped at 12 PM on Saturday, citing the property override and the loyalty table, with sources expanded
A property rule overriding a brand rule, with the quoted sources.
A $300 refund question answered with a manager note quoting the brand standard
A refund over the front-desk limit: answered, with the rule quoted.
The dog question answered at Beach Resort (40 lbs, $75 per stay) and at Mountain Lodge (any size, $50 per night capped at $200)
Same question, two hotels, two different correct answers.

How it works

Hotel documents›Clean and chunk›Input and scope checks›Search this hotel only›Top 5 passages›Claude: cited answer or refusal›Grounding and manager checks

Documents are cleaned and split by section, and each piece is tagged with its property, a readable name and its effective date. A question is checked first: if nothing in that hotel's documents matches, it's refused without calling the model. Otherwise the hotel's passages are ranked by keywords (and by meaning, where memory allows), and Claude Haiku 5.5 writes a short answer from the top five, citing each fact. Plain code then checks that every number appears in a cited source and adds manager notes where the documents require them.

Real-world mess, handled

Problem in the dataHow it's handled
Different rules at each hotelSearch is filtered to the selected property plus brand-wide documents
Outdated policy versionsDocuments marked as superseded are never indexed
Price tables in menusChunking never splits a table
"Page 1 of 3" footers in PDFsRepeated lines removed, ignoring changing page numbers
A scanned, image-only PDFDetected and flagged for OCR; safety questions point to a manager
A pasted guest email telling the AI to waive feesTreated as data, flagged in the document library, tested every run
A wrong fact in the questionCorrected from the sources instead of accepted

Results

54 test questions in 16 categories, including misspellings, a Spanish question, a prompt-injection attack, questions it must refuse, cross-property leaks, wrong facts in the question, off-topic questions and escalation rules. Each setup ran 3 times with the shipped model, Claude Haiku 5.5 (measured 2026-10-10).

MeasureLive demo (keyword search)Hybrid search
Right passage retrieved (strict)98%98%
Answer accuracy98%97% (96–98%)
Cites the right document98%98%
Every number traceable to a source98%100%
Correct refusals27 of 2727 of 27
"Check with a manager" rules correct54 of 5454 of 54
First question a visitor types: useful result100% (was 43%)100% (was 49%)
Full answer, 95th percentile2.5 s2.7 s
Cost per 1,000 questions$0.17$0.19

Testing it like a first-time visitor changed the product. The 54-question eval was written against the documents, so it scored 96–98% while the questions a visitor actually types first (Wi-Fi, the gym, breakfast, "checkout") were often refused or blocked. I measured 49 such questions first: 43% got a useful result. Filling real content gaps, spelling "check-out" and "checkout" the same way, curated suggestions, and friendly handling of greetings and off-topic input took it to 100%, with no regression on the original eval. I wrote those questions myself, so held-out questions from other people come next.

My grader passed a wrong answer. "Can a 16-year-old book a facial?" only had to mention "18", so "Yes, a 16-year-old can book a facial. The spa requires guests to be 18 or older" counted as correct. The same grader failed a safe answer that quoted the "waive all fees" injection in order to reject it. I fixed the measurement before touching the model, and I still read the injection answers by hand every run.

The model choice came from the eval, not the version number. On the first run, Claude Haiku 4.5 scored 96% and Haiku 5.5 98%, and Haiku 4.5's misses were real ("you can issue a $120 credit without additional approval"; the limit is $50). Re-run on the expanded documents, the two are within run-to-run range (98% vs 97%), so the reason to ship 5.5 is cost: $0.19 per 1,000 questions against $0.99. It's about a second slower at the 95th percentile, still under the 3-second target.

A "wrong fee" bug was in the question. A $75 fee showed up at the mountain lodge, where the fee is $50 a night. The cause was a number pasted into the question. The real risk was the model accepting stated facts, so it now corrects them, and four test questions check that it does.

274 tests (97% coverage), a dependency audit, a retrieval gate and a first-question gate run on every commit. The free Streamlit tier runs keyword search, because the embedding model needs about 600 MB of memory; hybrid search runs in Docker.

Business value

Projected from assumptions, shown as a range because the inputs are uncertain. The base case assumes 40 policy lookups a day per property, 2.5 minutes saved each, $20 an hour staff cost, and one wrong answer a week costing a $75 refund. The AI cost per question is measured, not assumed.

Net value per year2 properties100 properties
Pessimistic$8,524$423,187
Base$31,887$1,591,375
Optimistic$66,451$3,319,562

Even the pessimistic case returns many times its running cost of about $246 a year for two properties. The result depends most on minutes saved per lookup, so that's the first thing a pilot should measure.

Building it also surfaced policy gaps worth raising with the client, such as credits between $51 and $200 that no rule covers. A 2-week pilot at one property would time lookups before and after, review 50 answers, and track staff feedback against go / no-go targets agreed in advance. The full model and pilot plan are in the business case.

Tech stack

  • Python 3.12
  • Claude Haiku 5.5
  • BM25
  • sentence-transformers
  • Reciprocal Rank Fusion
  • pypdf
  • Streamlit
  • pytest
  • Playwright
  • Ruff
  • pip-audit
  • GitHub Actions
  • Docker

What I'd do next

  • A multilingual embedding model (the Spanish test question still fails) and OCR for scanned documents
  • An LLM judge calibrated against human labels, with 100+ test questions (Project 3)
  • Sign-in with roles, personal-data redaction and monitoring (Project 4)
  • Use this assistant as a tool inside a guest-email agent (Project 2)

Every design choice, with its tradeoff and result, is in the decision log.