Skip to content

Latest commit

Β 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

🍡 Matcha Café Recommendation System

A manually collected dataset and content-based recommendation system for exploring and recommending matcha drinks offered by cafΓ©s in Belgrade.

πŸ“Œ Project Overview

This project explores matcha drinks available in cafΓ©s across Belgrade. The dataset was manually collected from cafΓ© menus, websites, social media, Google Reviews, and food-delivery platforms such as Wolt.

The project aims to analyze the local matcha offering and develop a content- and aspect-based recommendation system. The recommender will suggest drinks based on characteristics such as sweetness, sentiment, ingredients, preparation style, and keywords extracted from customer reviews.

Installation and usage

git clone https://github.com/username/matcha-recommender.git
cd matcha-recommender
pip install -r requirements.txt

🎯 Objectives

  • Create a structured dataset of matcha drinks offered in Belgrade.
  • Analyze prices, ingredients, locations, and drink characteristics.
  • Extract sentiment, sweetness level, and keywords from customer reviews.
  • Explore spatial patterns in the local matcha offering.
  • Develop a content- and aspect-based recommendation system.

πŸ“₯ Data Collection

Data was collected manually from publicly available sources, including:

  • Official cafΓ© menus and websites
  • Instagram and other social media pages
  • Google Maps and Google Reviews
  • Customer-submitted photographs on Google Reviews
  • Wolt listings

Dataset Structure

Each row in the review-level dataset represents one customer review.

column_name description possible_values origin
cafe_id Unique numerical identifier for each cafΓ©. Integer (e.g. 1, 2, 3) Manually assigned
cafe_name Name of the cafΓ© as listed on Google Maps. Text (e.g. Wagokoro) Google Maps
location City district or neighborhood where the café is located. Text (e.g. Stari Grad, Vračar, Savski Venac) Google Maps / manually determined
latitude Geographic latitude coordinate of the cafΓ©. Decimal number (e.g. 44.81787914205736) Google Maps
longitude Geographic longitude coordinate of the cafΓ©. Decimal number (e.g. 20.472490255822375) Google Maps
is_chain Whether the cafΓ© belongs to a chain or franchise. TRUE / FALSE Manually determined
cafe_rating_google Overall cafΓ© rating displayed on Google Maps. Decimal from 1.0 to 5.0 (e.g. 4.8) Google Maps
num_reviews_cafe Total number of Google Maps reviews for the cafΓ©. Integer (e.g. 735) Google Maps
review_id Unique identifier for each review, created using cafe_id and a sequential number. Text (e.g. 1_01, 1_02, 2_01) Manually generated
drink_name Standardized English name of the drink. Text (e.g. matcha latte, cold strawberry matcha) Collected and manually standardized
flavor_type Syrup used to sweeten the matcha or another distinct flavor or aroma. Text (e.g. mango) CafΓ© menu, social media, Wolt, or manually inferred
liquid_type Type of milk or other liquid used in the drink. Text (e.g. almond, water) CafΓ© menu, reviews, images, or estimated
temperature Serving temperature of the drink. Text (e.g. hot, cold) CafΓ© menu, reviews, images, or estimated
drink_size_ml Estimated or stated volume of the drink in milliliters. Integer (e.g. 350, 360) CafΓ© menu, Wolt, social media, review images, or estimated
matcha_quality Matcha grade stated on the menu or visually inferred from images. Text (e.g. culinary, ceremonial) CafΓ© menu or visual estimate
price_rsd Price of the drink in Serbian dinars. Integer (e.g. 350, 400) CafΓ© menu, website, social media, or Wolt
price_per_100ml Calculated price per 100 ml: price_rsd / drink_size_ml * 100. Decimal number (e.g. 100.0, 114.28) Calculated
menu_descriptors Descriptive terms used by the cafΓ© to describe the drink. Comma-separated text (e.g. smooth, creamy, earthy) CafΓ© menu, website, social media, or Wolt
drink_rating Numerical rating given specifically to the drink when explicitly stated in a review. Integer from 1 to 5 Google Reviews
review_text Raw text of a Google Maps review mentioning the drink. Text Google Reviews; manually collected
sentiment_score Sentiment score extracted from review_text. Decimal from -1.0 to 1.0 / empty until processed NLP-derived
sweetness_level Sweetness level inferred from review_text. high / medium / low / unknown / empty until processed NLP-derived
keywords Key descriptive terms extracted from review_text. Comma-separated text / empty until processed NLP-derived

πŸ“ Project Structure

matcha/
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ processed/
β”‚   └── raw/ 
β”œβ”€β”€ images/
β”œβ”€β”€ notebooks/
β”‚   β”œβ”€β”€ 01_data_cleaning.ipynb
β”‚   β”œβ”€β”€ 02_eda.ipynb
β”‚   β”œβ”€β”€ 03_nlp_feature_extraction.ipynb
β”‚   └── 04_recommender_system.ipynb
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ feature_extraction.py
β”‚   └── recommender.py
β”œβ”€β”€ README.md
└── requirements.txt 

βš™οΈ Methodology

The project methodology consists of the following stages:

  1. Manual data collection and source documentation
  2. Data validation, cleaning, and preprocessing
  3. Exploratory data analysis
  4. NLP-based feature extraction:
    • Sentiment analysis using a pretrained sentiment model
    • Sweetness-level classification using a rule-based or model-based approach
    • Keyword extraction using TF-IDF, KeyBERT, or another suitable method
  5. Aggregation of review-level features into drink-level profiles
  6. Content- and aspect-based recommendation using cosine similarity or another suitable similarity measure
  7. Evaluation of the NLP pipeline and recommendation quality

πŸ“Š Exploratory Data Analysis

🚧 Status: In progress

This section will include key visualizations and findings related to cafΓ©s, drinks, prices, locations, and NLP-derived review features.

πŸ€– Recommendation Approach

Since the dataset does not contain user identifiers or interaction histories, collaborative filtering is not applicable. The system therefore uses a content- and aspect-based approach.

Each drink is represented using its structured attributes and review-derived features, including sweetness, sentiment, ingredients, and descriptive keywords. User preferences are compared with drink profiles to generate ranked recommendations.

πŸ“Š Results

This section will include:

  • Key findings from the exploratory data analysis
  • An example user query and the corresponding top five recommendations
  • Evaluation results for sentiment, sweetness-level, and keyword extraction
  • Evaluation of recommendation relevance and quality
  • A screenshot or demonstration of the final application

Limitations

  • It is impossible to determine the exact milk_type used for each drink. Many cafΓ©s only state on their menus that they offer a plant-based option, without specifying the exact type of milk. In some instances, the milk type could be derived from Google Reviews. Otherwise, cow’s milk was assumed to be the default option. For drinks such as matcha lemonade or yuzu matcha, the milk type was set to water, tonic water, orange juice or grapefruit juice for bumble.
  • It was also difficult to determine the exact drink temperature from Google Reviews. When users included a picture with their review, it was sometimes possible to determine whether the drink was hot or cold. Otherwise, the default value was set to cold, as iced matcha lattes are the most popular option.
  • Drink size was also impossible to determine with complete accuracy. When the drink_size_ml was stated on the menu, that value was used. Otherwise, it was estimated based on the cup size shown on social media, Instagram posts, or images submitted with Google Reviews. Wolt listings were also checked when available, as they sometimes included the drink size. Consequently, price_per_100ml is also an estimate in most cases.
  • Matcha type was estimated based on images posted on social media or included in Google Reviews. If the color appeared bright green, the matcha_type was classified as ceremonial grade, while duller or more yellow matcha was classified as culinary grade.
  • menu_descriptors were scarce. When available, they were derived from social media posts by local users and online menus found on Google, social media, Wolt, or the café’s website.
  • review_text and the other values described above were collected manually for this first version of the dataset.

🚧 Future Work

Task Description Status
Spatial analysis Split latitude_longitude into separate latitude and longitude variables. 🟒 Done
Data preparation Group all unspecified plant-based milk types under the plant category in milk_type. 🟒 Done
NLP feature extraction Derive and populate sentiment_score, sweetness_level, and keywords using NLP. 🟒 Done
Exploratory data analysis Conduct additional analyses and create informative visualizations. 🟑 In progress
Recommendation system Develop a content- and aspect-based recommender system. 🟑 In progress

πŸ” Reproducibility

To ensure that the project can be reproduced:

  • All required dependencies are listed in requirements.txt.
  • A fixed random seed is used where applicable.
  • The recommended order for running the notebooks is documented.
  • Raw and processed datasets are stored separately in the data/raw and data/processed directories.
  • NLP-derived columns can be regenerated by running the corresponding feature-extraction notebook.

βš–οΈ Ethical Considerations

The dataset contains publicly available Google Review text collected for research and educational purposes. No reviewer names, usernames, profile links, photographs, or other personal information are stored.

Review texts remain attributable to their original authors and source platform.

πŸ“„ License

This project is licensed under the MIT License.

πŸ‘€ Author

Ana Petrović
Master Engineer of Information Systems with a focus on data science and machine learning.

LinkedIn GitHub

About

🍡 Belgrade matcha reviews dataset with EDA, NLP feature extraction, and content-based recommendations.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages