A manually collected dataset and content-based recommendation system for exploring and recommending matcha drinks offered by cafΓ©s in Belgrade.
This project explores matcha drinks available in cafΓ©s across Belgrade. The dataset was manually collected from cafΓ© menus, websites, social media, Google Reviews, and food-delivery platforms such as Wolt.
The project aims to analyze the local matcha offering and develop a content- and aspect-based recommendation system. The recommender will suggest drinks based on characteristics such as sweetness, sentiment, ingredients, preparation style, and keywords extracted from customer reviews.
git clone https://github.com/username/matcha-recommender.git
cd matcha-recommender
pip install -r requirements.txt
- Create a structured dataset of matcha drinks offered in Belgrade.
- Analyze prices, ingredients, locations, and drink characteristics.
- Extract sentiment, sweetness level, and keywords from customer reviews.
- Explore spatial patterns in the local matcha offering.
- Develop a content- and aspect-based recommendation system.
Data was collected manually from publicly available sources, including:
- Official cafΓ© menus and websites
- Instagram and other social media pages
- Google Maps and Google Reviews
- Customer-submitted photographs on Google Reviews
- Wolt listings
Each row in the review-level dataset represents one customer review.
| column_name | description | possible_values | origin |
|---|---|---|---|
cafe_id |
Unique numerical identifier for each cafΓ©. | Integer (e.g. 1, 2, 3) |
Manually assigned |
cafe_name |
Name of the cafΓ© as listed on Google Maps. | Text (e.g. Wagokoro) |
Google Maps |
location |
City district or neighborhood where the cafΓ© is located. | Text (e.g. Stari Grad, VraΔar, Savski Venac) |
Google Maps / manually determined |
latitude |
Geographic latitude coordinate of the cafΓ©. | Decimal number (e.g. 44.81787914205736) |
Google Maps |
longitude |
Geographic longitude coordinate of the cafΓ©. | Decimal number (e.g. 20.472490255822375) |
Google Maps |
is_chain |
Whether the cafΓ© belongs to a chain or franchise. | TRUE / FALSE |
Manually determined |
cafe_rating_google |
Overall cafΓ© rating displayed on Google Maps. | Decimal from 1.0 to 5.0 (e.g. 4.8) |
Google Maps |
num_reviews_cafe |
Total number of Google Maps reviews for the cafΓ©. | Integer (e.g. 735) |
Google Maps |
review_id |
Unique identifier for each review, created using cafe_id and a sequential number. |
Text (e.g. 1_01, 1_02, 2_01) |
Manually generated |
drink_name |
Standardized English name of the drink. | Text (e.g. matcha latte, cold strawberry matcha) |
Collected and manually standardized |
flavor_type |
Syrup used to sweeten the matcha or another distinct flavor or aroma. | Text (e.g. mango) |
CafΓ© menu, social media, Wolt, or manually inferred |
liquid_type |
Type of milk or other liquid used in the drink. | Text (e.g. almond, water) |
CafΓ© menu, reviews, images, or estimated |
temperature |
Serving temperature of the drink. | Text (e.g. hot, cold) |
CafΓ© menu, reviews, images, or estimated |
drink_size_ml |
Estimated or stated volume of the drink in milliliters. | Integer (e.g. 350, 360) |
CafΓ© menu, Wolt, social media, review images, or estimated |
matcha_quality |
Matcha grade stated on the menu or visually inferred from images. | Text (e.g. culinary, ceremonial) |
CafΓ© menu or visual estimate |
price_rsd |
Price of the drink in Serbian dinars. | Integer (e.g. 350, 400) |
CafΓ© menu, website, social media, or Wolt |
price_per_100ml |
Calculated price per 100 ml: price_rsd / drink_size_ml * 100. |
Decimal number (e.g. 100.0, 114.28) |
Calculated |
menu_descriptors |
Descriptive terms used by the cafΓ© to describe the drink. | Comma-separated text (e.g. smooth, creamy, earthy) |
CafΓ© menu, website, social media, or Wolt |
drink_rating |
Numerical rating given specifically to the drink when explicitly stated in a review. | Integer from 1 to 5 |
Google Reviews |
review_text |
Raw text of a Google Maps review mentioning the drink. | Text | Google Reviews; manually collected |
sentiment_score |
Sentiment score extracted from review_text. |
Decimal from -1.0 to 1.0 / empty until processed |
NLP-derived |
sweetness_level |
Sweetness level inferred from review_text. |
high / medium / low / unknown / empty until processed |
NLP-derived |
keywords |
Key descriptive terms extracted from review_text. |
Comma-separated text / empty until processed | NLP-derived |
matcha/
βββ data/
β βββ processed/
β βββ raw/
βββ images/
βββ notebooks/
β βββ 01_data_cleaning.ipynb
β βββ 02_eda.ipynb
β βββ 03_nlp_feature_extraction.ipynb
β βββ 04_recommender_system.ipynb
βββ src/
β βββ feature_extraction.py
β βββ recommender.py
βββ README.md
βββ requirements.txt
The project methodology consists of the following stages:
- Manual data collection and source documentation
- Data validation, cleaning, and preprocessing
- Exploratory data analysis
- NLP-based feature extraction:
- Sentiment analysis using a pretrained sentiment model
- Sweetness-level classification using a rule-based or model-based approach
- Keyword extraction using TF-IDF, KeyBERT, or another suitable method
- Aggregation of review-level features into drink-level profiles
- Content- and aspect-based recommendation using cosine similarity or another suitable similarity measure
- Evaluation of the NLP pipeline and recommendation quality
π§ Status: In progress
This section will include key visualizations and findings related to cafΓ©s, drinks, prices, locations, and NLP-derived review features.
Since the dataset does not contain user identifiers or interaction histories, collaborative filtering is not applicable. The system therefore uses a content- and aspect-based approach.
Each drink is represented using its structured attributes and review-derived features, including sweetness, sentiment, ingredients, and descriptive keywords. User preferences are compared with drink profiles to generate ranked recommendations.
This section will include:
- Key findings from the exploratory data analysis
- An example user query and the corresponding top five recommendations
- Evaluation results for sentiment, sweetness-level, and keyword extraction
- Evaluation of recommendation relevance and quality
- A screenshot or demonstration of the final application
- It is impossible to determine the exact
milk_typeused for each drink. Many cafΓ©s only state on their menus that they offer a plant-based option, without specifying the exact type of milk. In some instances, the milk type could be derived from Google Reviews. Otherwise, cowβs milk was assumed to be the default option. For drinks such as matcha lemonade or yuzu matcha, the milk type was set to water, tonic water, orange juice or grapefruit juice for bumble. - It was also difficult to determine the exact drink
temperaturefrom Google Reviews. When users included a picture with their review, it was sometimes possible to determine whether the drink was hot or cold. Otherwise, the default value was set to cold, as iced matcha lattes are the most popular option. - Drink size was also impossible to determine with complete accuracy. When the
drink_size_mlwas stated on the menu, that value was used. Otherwise, it was estimated based on the cup size shown on social media, Instagram posts, or images submitted with Google Reviews. Wolt listings were also checked when available, as they sometimes included the drink size. Consequently, price_per_100ml is also an estimate in most cases. - Matcha type was estimated based on images posted on social media or included in Google Reviews. If the color appeared bright green, the
matcha_typewas classified as ceremonial grade, while duller or more yellow matcha was classified as culinary grade. menu_descriptorswere scarce. When available, they were derived from social media posts by local users and online menus found on Google, social media, Wolt, or the cafΓ©βs website.review_textand the other values described above were collected manually for this first version of the dataset.
| Task | Description | Status |
|---|---|---|
| Spatial analysis | Split latitude_longitude into separate latitude and longitude variables. |
π’ Done |
| Data preparation | Group all unspecified plant-based milk types under the plant category in milk_type. |
π’ Done |
| NLP feature extraction | Derive and populate sentiment_score, sweetness_level, and keywords using NLP. |
π’ Done |
| Exploratory data analysis | Conduct additional analyses and create informative visualizations. | π‘ In progress |
| Recommendation system | Develop a content- and aspect-based recommender system. | π‘ In progress |
To ensure that the project can be reproduced:
- All required dependencies are listed in
requirements.txt. - A fixed random seed is used where applicable.
- The recommended order for running the notebooks is documented.
- Raw and processed datasets are stored separately in the
data/rawanddata/processeddirectories. - NLP-derived columns can be regenerated by running the corresponding feature-extraction notebook.
The dataset contains publicly available Google Review text collected for research and educational purposes. No reviewer names, usernames, profile links, photographs, or other personal information are stored.
Review texts remain attributable to their original authors and source platform.
This project is licensed under the MIT License.
Ana PetroviΔ
Master Engineer of Information Systems with a focus on data science and machine learning.