hey there!

my name is

NAMITA SARDANA

As a data and AI professional bridging analysis, science, and engineering, I excel at designing, implementing, and fine-tuning data-driven algorithms for industry-specific applications, starting from the ground up with efficient extraction and preparation workflows to turn raw inputs into robust foundations. Every vertical comes with its own set of unique puzzles, and I love the challenge of adapting machine learning logic to fit those specific operational demands. My focus is always on building resilient pipelines that scale smoothly and adapt as needs evolve.

find my resumé here
See my work Portrait of Namita Sardana

experience

My experience sits mainly around data science, machine learning, and applied AI, with work spanning healthcare analytics, client retention, behavioural data, and AI-based distraction blocking. I've also worked on an earlier computer vision project involving image-based machine learning, while part-time work in hospitality has given me experience outside of the technical side of things. Across these experiences, I enjoy taking a problem from understanding the data and finding useful signals through to experimenting with models and turning the results into something practical.

Curious Cat

Data Science Intern · July 2025 – November 2025

Melbourne, Australia

At Curious Cat, the analytics division of Arrotex Pharmaceuticals, I worked on VetIQ, a data-driven platform for veterinary clinics. I was working with around 1 million records across 5 clinics, and a lot of my work came down to figuring out what we could actually learn from the way clients and their pets interacted with a clinic.

One of the things I led was the Opportunity Radar, which started from an idea I had during the internship. I didn't want to look at services independently and just report how often each one was being used. I was more interested in the relationships between them, what services tend to appear together, what patterns exist in the way clients use different services, and whether those patterns could point to an opportunity for a clinic.

I started by working through the transaction data and figuring out how to represent service usage in a way that made sense for the analysis. I used K-Means to explore groups of similar service usage behaviour and Apriori association rule mining to find relationships between services. That led to patterns such as Consultation leading to Medication and Preventive Care directing towards the next visit being for a Surgery.

Finding those relationships was only part of it. The analysis could produce a lot of different associations, so I worked through support, confidence and lift, along with the actual frequency of the services, to separate the more meaningful patterns from relationships that weren't particularly useful. I then worked on turning those into opportunity signals that could be shown through the VetIQ dashboard rather than leaving them as raw association rules.

I also contributed to the Client Retention Tracker, working alongside the person leading that piece. The idea was to see whether changes in a client's behaviour could give us an indication that they might not return to the clinic. I worked with transaction history to build recency, frequency and monetary value features, looking at how recently and how often someone visited and how their spending behaviour changed over time.

I was also involved in working out how to use the retention predictions in the Retention Tracker, rather than simply putting a model score in front of someone. The goal was to make changes in client behaviour easier to spot and understand through the dashboard.

What I liked about these two pieces was that they weren't just exercises in trying different models. With Opportunity Radar, I was trying to find patterns in the way services were being used that weren't obvious from looking at individual services. With the Retention Tracker, I was looking at how a client's behaviour changed over time. In both cases, a lot of the work was in getting the underlying data into a form that could actually answer those questions, deciding which patterns were worth paying attention to, and then figuring out how to present the results in a way that made sense outside of a notebook.

Focus Bear

AI Engineer Intern · November 2024 – March 2025

Melbourne, Australia

At Focus Bear, I worked on an AI-powered focus application designed to help adults with ADHD and autism manage digital distractions. My main project was a coarse website-filtering system that needed to quickly determine whether a website was relevant to what a user was trying to focus on.

The first version relied on regex and keyword-based rules. This made it fast and inexpensive, but it didn't handle context particularly well. A website could contain a relevant keyword without actually being useful for the user's goal, while another relevant website might use completely different wording. I therefore moved the approach towards understanding the semantic relationship between the user's intention and the website content.

I experimented with BERT-based representations to capture context and meaning rather than relying solely on lexical matches. This made it possible to treat the user's goal and the website information as related inputs to the classification process, allowing the system to distinguish between superficially similar content with different purposes.

The final approach combined deterministic rules with machine-learning signals. This hybrid design was important because the problem wasn't just about getting the highest possible classification score. The filter had to make decisions quickly, reduce unnecessary GPT-4o API calls, control inference costs, and still behave sensibly when presented with websites that hadn't appeared in the training or testing data.

A large part of the work involved evaluating the trade-offs between precision, recall, latency, and cost. I tested different thresholds and approaches, investigated false positives and false negatives, and worked through edge cases where the correct behaviour wasn't immediately obvious. The resulting system improved classification accuracy by approximately 5.5% while reducing the reliance on a general-purpose LLM for every filtering decision.

This project gave me practical experience taking an AI feature from a simple heuristic implementation through to a more context-aware NLP system, while dealing with the constraints that come with putting machine learning into an actual product, particularly inference speed, API cost, model behaviour, edge cases, and reliability on unseen inputs.

projects

These are some of the projects I have been working on recently, covering areas like sports analytics, applied AI, and environmental data. They are quite different in what they are trying to solve, but most of them have involved the same general process for me: understanding the problem first, figuring out what the data can actually tell me, experimenting with different approaches, and then turning the results into something that is useful beyond just the model itself.

T20 Match Engine

Sports Analytics Framework

The T20 Match Engine started from an interest in looking at cricket through something more meaningful than traditional scorecard statistics. I spent some time researching how professional cricket analytics platforms such as CricViz approach performance evaluation and used that as a starting point for thinking about what I could build and improve upon myself. A batter scoring 40 runs, for example, doesn't tell us much on its own about whether those runs came when the team needed them most, while a fielder's contribution can be difficult to capture through standard statistics at all. I wanted to build a system that could take the match situation into account when evaluating what actually happened on the field.

I approached the problem by treating the two innings differently rather than trying to force them into one model. For the first innings, I built an XGBoost regression model to estimate the expected final score from information available at each point in the innings, including match state, venue conditions, scoring rate, and the phase of the game. This became the basis for Expected Score (eScore) and allowed me to compare a batter's actual contribution against what would normally have been expected from that situation.

For the second innings, the question changes from "how many runs will this team score?" to "how likely are they to win from here?". I therefore built a separate XGBoost classifier that continuously estimates Win Probability (WP) throughout a chase. From this, I could look at how individual deliveries changed the outcome of the match and use that to calculate Win Probability Added (WPA) for players.

Once those two models were working, I used their outputs to build some of the other metrics I was interested in. These include Runs Above Expectation (RAE) for batters, Expected Runs Saved (ERS) for fielders, and a dynamic Leverage Index (LI) that measures how much pressure a particular delivery carries based on the state of the match. I also experimented with a Clutch Score for weighting performances in high-pressure situations and a Pressure Index to capture the cumulative effect of things like consecutive dot balls.

A big part of the project has been making sure these metrics actually lead to useful analysis rather than just producing more numbers. The framework can be used to compare players in context, identify game-defining deliveries, and break down team performance across the powerplay, middle overs, and death overs. I have also been working towards presenting the outputs through a dashboard so that the analysis can move naturally from first-innings target projection to live win probability during the chase.

What I find particularly interesting about this project is that the underlying idea isn't limited to cricket. The same approach can be adapted to other sports where the game can be viewed as a sequence of phases or possessions. In basketball, for example, the same framework could be extended towards expected possession value, live win probability, leverage, and late-game decision making.

Eunoia

RAG Chatbot for Mental Health Support

Eunoia came out of a university group project where we wanted to explore how conversational AI could be used to support mental wellbeing without trying to replace professional or clinical support. The main challenge wasn't simply getting a language model to generate a response. We wanted to build something that could respond in a way that was grounded in reliable information while also knowing when it shouldn't try to answer.

We built the system around a Retrieval Augmented Generation (RAG) pipeline running locally through Ollama. I worked with nomic-embed-text to generate embeddings for the knowledge base and FAISS to index and retrieve relevant information. Mistral was then used to generate responses based on the retrieved context rather than relying entirely on the model's own knowledge.

The knowledge base combined curated psychological guides and coping resources with a questions.csv dataset containing conversational examples and human-aligned responses. This gave us two different types of information to work with: factual resources that could ground the model and examples that helped it understand the sort of language people might actually use when talking about their wellbeing.

One of the things we found early on was that straightforward keyword matching wasn't enough for the sort of queries we wanted Eunoia to handle. Someone saying they feel "burnt out but not sad", for example, needs a response based on the meaning of the query rather than simply the words it contains. Using semantic embeddings allowed the retrieval layer to find information based on context and meaning, and our approach improved contextual retrieval precision by more than 5% compared with the baseline we started with.

We also built a feedback loop around the system so that responses could be evaluated and used to improve the retrieval process. Through this refinement, retrieval accuracy improved by around 8%, while helpfulness improved by approximately 2%. This was useful because it gave us a way to evaluate the system based on how people actually experienced the responses rather than only looking at technical retrieval metrics.

Safety was another important part of the design. If the system isn't confident enough in the information it has retrieved, it doesn't simply generate an answer and hope for the best. Instead, a safety fallback redirects the user towards trusted support services such as Lifeline or Headspace. For me, this was one of the more important parts of the project because it showed how the design of an AI system has to account for what happens when the model isn't reliable, rather than only focusing on what happens when it works.

CarbonWise

Personal Environmental Impact Tracker

CarbonWise started from a fairly simple question: how much carbon do we actually produce just from getting around every day? I originally planned to use my Google Maps Takeout data to reconstruct my own travel history, but because location tracking had been turned off, I didn't have enough real historical data to work with. Instead of abandoning the idea, I built a synthetic 23-year travel history that represented different stages of my life, from childhood trips and school buses through to university commutes and part-time work.

I used that data to build a dashboard that breaks down the resulting travel behaviour and emissions. It tracks total CO₂ emissions, distance travelled, transport modes, flight emissions, and monthly emissions, while also comparing actual emissions against a personal monthly target. The idea was to make the numbers easier to understand by showing not just a total footprint, but where that footprint was actually coming from.

I also wanted the project to go a step further than simply telling someone that their emissions are high. I built a decision tree model that looks at things like trip distance, the type of journey, and its potential environmental impact to determine when alternatives such as walking, cycling, or public transport might make sense. This means the dashboard can move from describing past behaviour to suggesting what could be changed.

Another part of the project was making the system usable beyond my own simulated travel history. Users can add their own trips and have the same calculations and recommendations applied to them. This meant thinking about the project less as a one-off analysis and more as a small data application, where new inputs need to pass through the same processing and modelling pipeline before being turned into something meaningful.

What I liked about CarbonWise was that it gave me a chance to work on a project where the model is only one part of the solution. The more interesting challenge was connecting data generation, emissions calculations, machine learning, and visualisation so that someone could actually use the results to understand their own behaviour and think about possible alternatives.

technologies

There isn't really one stack I stick to. Most of my work sits somewhere between data, machine learning, AI, and engineering, so the tools I use tend to change depending on what I'm building and what the problem calls for.

I'm most comfortable working with Python and the wider data science ecosystem, but I've also worked across R, SQL, cloud platforms, distributed computing, machine learning, NLP, and web technologies. Some tools I use regularly, while others are technologies I've worked with on particular projects and continue to build on.

languages

PYTHON

My main language for data analysis, machine learning, AI, automation, and building data-driven applications.

R

Used for statistical analysis, data exploration, modelling, and visualisation.

SQL / MySQL

Querying, joining, cleaning, transforming, and analysing structured data and relational databases.

PL/SQL

Working with relational data through database-side logic, procedures, and queries.

MONGODB

Working with document-based data and less rigid NoSQL structures.

WEB

HTML, CSS, and JavaScript for web interfaces and lightweight applications.

data science & machine learning

PANDAS & NUMPY

My usual starting point for cleaning, transforming, exploring, and preparing data for modelling.

SCIKIT-LEARN

My go-to toolkit for preprocessing, feature engineering, model training, and evaluation.

TENSORFLOW & PYTORCH

Used for deep learning and neural network-based projects, from experimentation through to model development.

XGBOOST

Particularly useful for structured-data problems where gradient-boosted models are a good fit.

MLFLOW & WEIGHTS & BIASES

Used around experimentation to track runs, models, metrics, and development decisions.

ai, nlp & retrieval

HUGGING FACE

Pretrained models, tokenisers, embeddings, text classification, question answering, and generative AI experiments.

LANGCHAIN

Connecting language models with documents, retrievers, vector stores, prompts, and other application components.

QDRANT

Vector search with embeddings, metadata, filtering, and retrieval workflows.

WEAVIATE

Semantic retrieval combining vector similarity with metadata and structured information.

Ollama

Running language models locally for experimenting with models, prompts, and retrieval workflows.

Mistral

Open-weight language models for generation, summarisation, and instruction-based tasks.

RAGAS

Evaluating retrieval and generated responses in knowledge-based AI applications.

cloud & mlops

AMAZON WEB SERVICES

AWS environments for data and machine learning workloads, including distributed analytics with EMR.

GOOGLE CLOUD

Used for machine learning workloads, experimentation, and cloud-based solutions.

MICROSOFT AZURE

Familiarity with Azure Storage, Azure AI services, and deploying applications and ML workloads.

DOCKER

Containerisation basics for consistent and reproducible application environments.

databases & search

MYSQL

Relational database work, from querying and transforming data to supporting analytical workflows.

MONGODB

Document-based storage for applications using a flexible NoSQL data model.

PLANETSCALE

MySQL-compatible cloud databases in application environments.

ELASTICSEARCH

Indexing and searching large collections of structured and unstructured information.

FIERBASE

Cloud services and application infrastructure for web and data-driven applications.

visualisation

TABLEAU & POWER BI

Building dashboards and interactive reports that make data easier to explore and communicate.

MATPLOTLIB & SEABORN

Python-based visualisation for exploring datasets, understanding patterns, and communicating results.

GGPLOT2

R-based visualisation for statistical analysis and presenting data clearly.

apis, applications & automation

FLASK & FASTAPI

Building APIs and lightweight services that expose data, machine learning models, and AI functionality.

GRAPHQL

Working with APIs where applications need more flexible ways to request and consume data.

N8N

Automating workflows and connecting applications, services, and data sources.

RAILWAY

Deploying and hosting applications and containerised services without managing the underlying infrastructure.

development & collaboration

GIT & GITHUB

Version control, collaboration, code management, and keeping development work organised.

TRELLO

Keeping track of tasks, project work, and the moving pieces behind larger projects.

OVERLEAF

Collaborative technical writing, research documentation, and LaTeX-based projects.