Skip to content

About

A Generative AI framework for Digital Ethnography: Decoding local authenticity and cultural signals in Vietnamese gastronomy reviews.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

15 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Capstone-AI-Humanities

A Generative AI framework for Digital Ethnography: Decoding local authenticity and cultural signals in Vietnamese gastronomy reviews.

The Authentic Gaze: Decoding Gastronomic Heritage

Course: 2026S 136031-1 GenAI for Humanists Author: Nguyen Ngoc Huyen Tran Matrikelnummer: 11922332 Status: Delivery 2 — Complete


1. Personal Motivation & Context

This project is inspired by my recent journey through Vietnam. As a member of the Vietnamese diaspora born in the West, my goal was to reconnect with my heritage by seeking out "local-approved" culinary experiences.

During this trip, I realized that platforms like Google Maps often prioritize high-volume, tourist-rated locations. I found myself manually filtering through hundreds of reviews, looking for Vietnamese names and non-translated "Teencode" to find the "real" food. This project uses Generative AI to study that filtering process rather than simply automate it — see Section 3 for why this distinction matters.

2. Project Objectives

  • Deconstruct the "Tourist Bubble": Examine how review platforms surface certain linguistic and cultural signals over others, and what gets systematically pushed down.
  • Bridge the Diaspora Gap: Build a prototype that helps "Việt Kiều" (Overseas Vietnamese) read reviews through a more locally-grounded lens.
  • Demonstrate a Re-Ranking Method: Produce a working demonstration — not a deployed product — showing how an alternative, locally-weighted ranking could look for a small, hand-picked sample of venues.

3. Relevance to Humanistic Inquiry & Literature Review

This project sits at the intersection of Digital Ethnography, Tourism Studies, and Gastropolitics. Rather than treating "authenticity" as a fact to be detected, the project treats it as a discourse — something produced through language, platform design, and the positionality of the reviewer.

  • John Urry — The Tourist Gaze (1990): Tourists see places through socially constructed expectations of what is worth looking at. Review platforms can be read as infrastructure that organises and reproduces this gaze, surfacing what conforms to outsider expectations of Vietnamese food.
  • Dean MacCannell — Staged Authenticity (1973): MacCannell's concept describes how spaces are deliberately presented to appear authentic to visitors while remaining a performance. This is directly relevant to vendors who may perform "localness" for a tourist audience.
  • Arjun Appadurai — Gastropolitics (1981): Appadurai's framing of food as a site where social and political relations are negotiated supports treating restaurant reviews as a political text — claims of who has the right to evaluate, define, or "own" a cuisine.
  • Lugosi & Bell — Platform-mediated food culture: Their work on hospitality and digital platforms is relevant to how Google Maps reviews are shaped by platform incentives (star ratings, review-volume sorting, language defaults).

4. The Central Research Question

What signals do reviewers code as "authentic," and whose interests does that coding serve?

Working sub-questions:

  • Do Vietnamese-language and English-language reviews of the same venue diverge in what they praise?
  • Does the platform's default sort correlate with tourist-coded language over locally-coded language?
  • What do reviewers who use insider language have in common, and what do they implicitly exclude?

5. AI Techniques Employed

  • Zero-Shot Classification: facebook/bart-large-mnli — a local model requiring no API key, run entirely within Google Colab.
  • Sentiment Weighting: Categorizing what reviews praise (broth depth, traditional technique) versus tourist-infrastructure markers (English-speaking staff, décor, air-conditioning).
  • Zero-Shot Persona Tagging: Classifying reviewers as "Cultural Insider" vs "Casual Traveler" based on the six-feature checklist in Section 6.

6. The Classifier: Operationalizing Cultural Insider vs. Casual Traveler

The classifier was driven by six explicit linguistic features:

# Feature Signal
F1 Review written in Vietnamese Language choice
F2 Uses regional slang or "Teencode" Colloquial register
F3 Mentions specific techniques or ingredients Content depth
F4 Compares to other local venues or family standard Local frame of reference
F5 Does NOT focus on tourist infrastructure as main criteria Absence of outsider lens
F6 Uses regional dialect markers Geographic specificity

Rule: A minimum of 2 features must be present to classify as Cultural Insider. Language alone (F1) is not sufficient.

7. Dataset & Validation

  • 45 reviews manually collected across 3 Cơm Tấm venues in Saigon
  • All reviews hand-coded by the researcher before running the model
  • Model classifications compared against the hand-coded ground truth
Metric Result
Total reviews 45
Percent agreement 64.4%
False Positives (AI over-called Insider) documented in error log
False Negatives (AI missed Insider) documented in error log

8. Delivery 2: Results & Findings

8.1 The Re-Ranking Output

The Streamlit dashboard re-ranks the three venues by an "Authenticity Score" — a weighted combination of human insider ratio, AI insider ratio, and a tourist-keyword penalty — rather than by platform star rating or review volume.

Rank Venue Authenticity Score
🥇 1 Com Tam @2 Sai Gon 76.0%
🥈 2 Cơm tấm Đề Thám ~70%
🥉 3 Cơm Tấm Tốp Mỡ - Authentic Saigon Broken Rice ~65%

This re-ranking diverges from the platform's default sort, which orders by review count and star average — demonstrating that the two ranking logics (volume-based vs. cultural-signal-based) produce different results and surface different venues.

8.2 Error Analysis: Three Failure Modes

The 35.6% of cases where the AI disagreed with the human coder cluster into three identifiable patterns:

1. The Foreign Language Trap The model could not distinguish Vietnamese from Japanese or Korean. Any non-Latin script triggered an implicit "cultural insider" assumption. This is a direct failure of F1 over-dominating F3/F4 — the model confused script with cultural competency.

2. The English-Menu-as-Bonus Error The model read tourist-infrastructure language ("great English from our waiter," "easy to pay by card") as a positive amenity rather than an outsider-orientation signal, and counted it in favor of a positive experience rather than as a marker of tourist-coded framing.

3. The Enthusiasm Trap / Staged Authenticity in the Model Warm, enthusiastic tone — even in English, even focused on décor — was sometimes read by the model as insider affection for the venue. It classified the performance of local warmth rather than its substance. The "free bananas from the artificial banana tree" review is the clearest example: vivid, affectionate, specific — and entirely tourist-oriented.

9. Conclusion

The 64.4% agreement rate is not a disappointment. It is the finding.

The experiment revealed that the AI partially followed the rules — it learned to flag food-specific language and non-English text as insider signals — but it could not replicate the tacit cultural knowledge that a diaspora researcher brings to the same task. The gap between what a six-feature checklist can formalise and what cultural competence actually does is precisely where the most interesting analytical work lies.

This is MacCannell's staged authenticity playing out inside the model itself: the AI was fooled by the performance of local knowledge, just as tourists are fooled by the performance of local atmosphere. It classified the staging, not the substance.

What a human eye catches — and the model misses — is implicit knowledge: the way a Vietnamese speaker recognises that "bonus points for great English" signals the reviewer treats language as an amenity rather than a given, or that praising a "charming artificial banana tree" is performing Saigon nostalgia for an outside gaze rather than describing a familiar neighbourhood spot.

The conclusion is therefore not "the AI followed my rules at 64%." It is: formalising cultural insider knowledge into a six-feature checklist is itself an act of reduction, and the 36% of cases where the AI failed mark exactly where that reduction loses something irreplaceable. This connects directly to Appadurai: the right to evaluate a cuisine is not just a linguistic pattern. It is an embodied, relational, historically-situated claim — and that is something a zero-shot classifier cannot fully capture. The tool demonstrates its own limits, and those limits are the argument.

10. Ethics

  • Data collection: Reviews were manually collected in compliance with Google Maps' Terms of Service. No automated scraping was used. The dataset is small (45 reviews) and treated as a demonstration sample, not a representative corpus.
  • Reviewer privacy: All reviewer usernames have been anonymized in any public-facing output. No profile photos or identifying information are included.
  • The politics of labeling: Tagging individuals as "Cultural Insider" vs "Casual Traveler" is itself an act of gatekeeping. The classifier's output is framed as the model's reading of linguistic signals, not a judgment on any individual's identity or right to an opinion about Vietnamese food.
  • Positionality: As a member of the diaspora, the researcher occupies an ambiguous position — neither a "local" by current residence nor a "tourist" by heritage. The design choices in the feature checklist (which signals count as "authentic") reflect this particular diaspora vantage point and are not presented as neutral or universal.

11. Repository Structure

/colab/
    authentic_gaze_classifier.ipynb   ← classification pipeline
/streamlit/
    app.py                            ← dashboard code
/data/
    reviews.json                      ← raw reviews + manual labels
    classified_reviews.json           ← reviews + AI labels + agreement flag
README.md

About

A Generative AI framework for Digital Ethnography: Decoding local authenticity and cultural signals in Vietnamese gastronomy reviews.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages