Skip to content
View parhamkhoshsolat's full-sized avatar

Highlights

  • Pro

Block or report parhamkhoshsolat

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
parhamkhoshsolat/README.md

Parham Khosh Solat

Data Scientist · ML Engineer Three years building the data and recommendation layer of a live influencer marketing platform. MSc Data Science · University of Naples Federico II · Graduating October 2026

Portfolio · LinkedIn · HuggingFace · parhamkhoshsolat@gmail.com


What I do

I ship machine learning end to end. For three years that meant production work at Sakoudar, an influencer marketing platform, where I built the ingestion, the scoring and the recommendation layer for a product that is still running. Alongside it, several projects are live as interactive demos on HuggingFace Spaces. My MSc thesis is on transparency in human-robot reinforcement learning.


Experience

Data Scientist · Sakoudar · 2022 to 2025 · full time Influencer campaign management platform across LinkedIn, X and Instagram. Joined in the company's first year.

  • Built the weekly ingestion across the LinkedIn, X and Instagram APIs, turning one-off snapshots into a per-influencer time series.
  • Designed the scoring that made influencers comparable, weighting impressions so a daily poster did not simply outrank a weekly one, and so categories with very different audience sizes could be judged fairly.
  • Built the recommendation engine: a customer brief in (budget, platforms, vertical, headcount, and the company's own mission and voice), a ranked shortlist out, matched partly on how closely a creator's language resembled the way the company talks about itself.
  • Built the dashboard the founders used to check every score against what they personally knew about each influencer, and the pipeline that pushed data live only after they signed off.

Pinned projects

Project Stack What it does
Shielded LGV Routing PyTorch, Bayes-DQN, CBS, PIBT, WebSocket Six warehouse vehicles routed by nine coordination methods, classical planners against reinforcement learning, all under a PIBT collision shield that holds collisions at zero so methods compete on throughput alone. The classical planners won. Live demo
RL-Restore PyTorch, DQN+LSTM, Real-ESRGAN, FastAPI, React Reimplemented a CVPR 2018 paper: a DQN+LSTM agent repairs photos by chaining 12 specialist CNNs. Extended with baselines, the paper's unreleased joint fine-tuning, and a live web app that shows every step. Live demo
Florence-2 VQA PyTorch, HuggingFace Transformers, Streamlit Full fine-tune of Microsoft Florence-2 (230M params, every parameter trainable) for visual question answering, on 60,000 image-question pairs from VQA v2.0. Live demo
Retail Geospatial Analytics GeoPandas, Folium, Python, SQL Industry analytics challenge with Fater S.p.A. (Procter & Gamble × Angelini Industries). Joined proprietary sales data with public census data at district level, and designed a per-capita metric so districts of different sizes compared fairly. Presented the findings to company leadership; the four-person team was recognised for the project.
Pest Population Forecasting Scikit-learn, XGBoost, LightGBM, Streamlit Benchmarked six regression and five classification models on noisy multi-source sensor data. Random Forest topped both, catching every outbreak day in the held-out set at 0.92 AUC and roughly 50 percent precision, with the threshold tuned to favour recall. Live demo
Stock Clustering Pipeline Apache Kafka, PySpark, scikit-learn Kafka and Zookeeper cluster ingesting six months of daily prices across 55 per-ticker topics, plus a separate PySpark MLlib K-means and PCA stage on a ticker's price series.
OULAD Time-Series Forecasting TensorFlow / Keras, Statsmodels, Prophet Benchmarked SARIMA, ARIMAX, Prophet and a custom 1D CNN on student interaction data. The CNN and Prophet came out closest; only the CNN was scored on a held-out split, so it is not a like-for-like comparison.

Skills

Languages: Python, SQL, JavaScript, Bash ML / Deep Learning: Recommender systems, text similarity, sentiment analysis, PyTorch, HuggingFace Transformers, Scikit-learn, XGBoost, LightGBM, TensorFlow / Keras, Random Forest, LSTM, GRU, CNN, Reinforcement Learning, Fine-tuning, Transfer Learning Data Engineering: Apache Kafka, PySpark, ETL, REST APIs, GraphQL, MySQL Visualisation & BI: Power BI, Plotly, Seaborn, Matplotlib, GeoPandas, Folium, Streamlit MLOps: Docker, FastAPI, HuggingFace Spaces, Git / GitHub workflows, Colab GPU training


Credentials

  • MSc Data Science at Federico II (in progress, expected Oct 2026) · weighted average 28.67/30 · four exams at 30 e lode
  • 5G Academy · Federico II with Nokia, TIM, and PagoPA (currently attending)
  • Apple Foundation Program · Federico II × Apple Developer Academy (Jan 2025)
  • BSc Information Technology Engineering · Amol University (2017)

Now

Working on my MSc thesis: research in human-robot interaction and reinforcement learning, extending published work from Federico II.

Open to Data Analyst, Data Scientist, ML Engineer, or AI Engineer roles starting now. Onsite Naples or remote across the EU.

📧 parhamkhoshsolat@gmail.com · Portfolio · LinkedIn

Pinned Loading

  1. florence2-vqa florence2-vqa Public

    Full fine-tune of Florence-2 (230M) for visual question answering. Live on HuggingFace Spaces.

    Jupyter Notebook

  2. TalentSonar TalentSonar Public

    Unfinished 24-hour hackathon prototype. Not a working product.

    Python

  3. retail-geospatial-analytics retail-geospatial-analytics Public

    Geospatial retail analytics for a Fater industry challenge. Public version runs on synthetic data.

    Jupyter Notebook

  4. pest-population-forecasting pest-population-forecasting Public

    Regression and classification pipelines for pest risk prediction. Deployed on HuggingFace Spaces.

    Jupyter Notebook

  5. stock-clustering-pipeline stock-clustering-pipeline Public

    Kafka ingestion across 55 per-ticker topics, plus a PySpark K-means and PCA clustering stage.

    Jupyter Notebook

  6. time-series-OULAD time-series-OULAD Public

    Time-series forecasting benchmark: SARIMA, ARIMAX, Prophet, CNN on student interaction data

    Jupyter Notebook 1