Senjuti Bala

Senjuti Bala

I am a data scientist with roots in research, consulting, and software development. I enjoy working at the intersection of data, systems, and behavioral science. I trained as an engineer and spent two years in technical consulting for telecom, fintech, and consumer goods clients in Bangladesh, primarily on product and financial modelling. I am most at home figuring out how patterns actually work. Agents are my friends, and a simulation is my playground. I also co-authored a book chapter in Fair Data Fair Africa Fair World, where I built a knowledge graph from the testimonies of trafficking survivors. My MSc. thesis introduced a two-layer agent-based model of climate negotiations among three types of countries. I am currently extending it so the agents can learn to make better choices.

I am based in Leiden, the Netherlands, and currently looking for opportunities in data and applied AI. I received my master’s degree from Leiden University in computer science. I previously studied information technology at the National Institute of Technology, Durgapur.

Skills: agent-based modelling, statistical modelling, RAG, Python (Pandas, NumPy, Scikit-learn, Matplotlib, Statsmodels, Plotly, Mesa), PyTorch, SQL, FAIR data, OWL/RDF ontology design, SPARQL, Power BI, and Streamlit. Fluent in English and Bengali, and a beginner in Dutch.

Experience

Projects

MSc thesis, Leiden University (LIACS), 2025

Two-Layer Agent-Based Model for Climate Cooperation

A two-layer agent-based model that couples national climate-policy agents (developed, developing, and vulnerable country types) with consumer-level green technology adoption. The model tests fairness mechanisms, international aid flows, and bounded rationality across 7 scenarios over a 100-year horizon, run across 100 seeds. The key finding is that access matters more than awareness. Code is available on GitHub.

PythonMesa ABMAgent-Based ModellingPandas

Co-authored chapter, Fair Data Fair Africa Fair World, 2025

FAIR Data Implementation for Analysis of Research Data in Human Trafficking and Migration

I co-led the FAIRification of 126 ethnographic interviews with trafficking survivors across Libya, Sudan, Ethiopia, and Europe. I built a custom OWL ontology, converted records to RDF and Turtle using rdflib, and loaded them into AllegroGraph for SPARQL querying. I also built a Streamlit dashboard and Gephi network graphs for non-technical researchers, under full DPIA and GDPR compliance. DOI: 10.5281/zenodo.15383037

FAIR PrinciplesRDF/OWLSPARQLStreamlitNetwork Analysis

Independent project

Job Database: NL Job Search Tool with Entity Resolution

Originally built for personal use, after getting tired of manually checking which companies actually sponsor work visas for non-EU applicants. The Dutch sponsor register lists companies by legal name, like Booking.com International B.V., while job listings use brand names, like Booking.com, so simple text matching fails. This tool fixes that with embeddings and vector search in BigQuery, matching company names by meaning instead of exact text. It also ranks companies by how well they match a pasted CV, using the same embedding approach. Backend built in Flask.

PythonFlaskBigQueryVector SearchEntity Resolution

Independent project

Crop Yield Prediction for SDG2

A machine learning model that predicts crop yields from agricultural and climate data, built in support of UN Sustainable Development Goal 2, Zero Hunger.

PythonScikit-learnRegression

Independent project

Bias Detection in News Sentences

A text classification model that detects bias at the sentence level in news articles.

PythonNLPText Classification

Independent project

Neural Reranking for MS MARCO

A neural reranking model for passage retrieval on the MS MARCO dataset.

PythonPyTorchInformation RetrievalTransformers

Independent project

AI Grandma: A RAG-Based Folklore Storyteller

Built to explore different cultures through their myths and folk stories. A retrieval-augmented storyteller grounded in real folk tales instead of generic prompts. Chunks and embeds the tales into a FAISS vector store, retrieves the closest match to a chosen theme, generates a retelling with a local LLM, and narrates it aloud through a Streamlit app. Runs fully offline on open-weight models.

PythonRAGFAISSStreamlitLocal LLM

Education

Leiden University

MSc Computer Science, Data Science Specialisation

2023 to 2025

National Institute of Technology, Durgapur

BTech, Information Technology

2016 to 2020

Notes

Outstanding Agility Award, twice, RedDot Digital. Awarded for consistent delivery across KPIs, client impact, and stakeholder communication.

ICCR Full Scholarship: a government-funded, four-year BTech. JENESYS Exchange Fellowship: a fully funded exchange to Japan.

Volunteers with Green Kitchen Leiden on food waste reduction.

Hobbies include long-distance trekking, cycling, and films.

save the bees
Open to work