Abstract dark render of a glowing disc wrapped in a data mesh
Applied Data Scientist · Iż-Żurrieq, Malta

ArjunKhanna

Made for stakeholders. Built for scale.

Pipelines, predictions, and agentic AI systems. I turn multi-dimensional data into decisions — and build the systems that keep doing it without me.

Not just models. 

End-to-end systems that keep predicting without me.

PythonSQL / T-SQLScikit-LearnPandasNumPyPyTorchXGBoostFastAPIMLflowPower BIC#Agentic AIPythonSQL / T-SQLScikit-LearnPandasNumPyPyTorchXGBoostFastAPIMLflowPower BIC#Agentic AI

About me

01 / About

Dedicated Applied Data Sciences candidate combining rigorous academic training with practical experience in machine learning, data engineering, and AI system architecture. I design end-to-end predictive pipelines, automate enterprise data workflows, and run robust statistical analysis, including anomaly detection and A/B testing. 

Proficient in Python and SQL/T-SQL, with a proven ability to translate multi-dimensional datasets into strategic business intelligence, optimizing operational efficiency and mitigating risk across technology, consulting, and manufacturing.

End-to-end ML pipelinesETL & data engineeringAnomaly detection & riskAgentic AI systems
Currently
BSc (Honours) in Applied Data Sciences, MCAST
Based in
Iż-Żurrieq, Malta
Born
Melbourne, Australia
Nationality
Indian
Languages
Hindi (native) · English (C2)
0h

Weekly hours reclaimed by ETL automation at GRID

From a month of manual CSV joins to a scheduled pipeline.

0+

Enterprise APIs wired into the Flying Sapphire agent

Google Workspace, Slack, and logistics tooling — all on-device.

0%

Score on the A/B testing machine-learning assignment

Wine quality, hypothesis-tested and peer-graded.

0

End-to-end ML pipelines shipped

From raw source to decision and dashboard.

What I bring

Three practice areas — 02 / Offerings

Tier 011 stack — Pandas → SQL → schedule

Data & ETL Engineering

I connect, clean, and automate.

  • —Automated ETL pipelines (CSV → warehouse, unattended)
  • —SQL / T-SQL query architecture & optimization
  • —Data governance, quality and structural integrity
  • —Python tooling that removes the manual step entirely

Teams drowning in weekly manual data prep.

Tier 022 stacks — Scikit-Learn → PyTorch → LLMs

Machine Learning & AI

I build predictions that hold up.

  • —Predictive modeling, classification & forecasting
  • —Anomaly detection for fraud and risk
  • —Hyperparameter tuning and MLOps tracking (MLflow)
  • —Agentic AI systems driven by local LLMs

Risk, forecasting, and decision automation.

Tier 033 stacks — Analytics → A/B → Power BI

Data Intelligence

I make data legible.

  • —Exploratory analysis and strategic business intelligence
  • —A/B testing and hypothesis-driven decisioning
  • —Real-time KPI dashboards stakeholders actually read
  • —Data storytelling that turns findings into decisions

Leaders who want decisions, not spreadsheets.

Where I've worked

03 / Experience

GRID

Data Scientist

03/2024 – 12/2024

Amritsar, India

ETLSQLDashboardingStorytelling
  1. 01

    Exploratory Data Analysis & Wrangling

    Executed comprehensive data cleaning, wrangling, and re-shaping on raw client datasets, establishing strict data-quality gates before any statistical work touched the data.

    Cleaner inputs → trustworthy outputs

  2. 02

    Database Query Optimization

    Architected and optimized SQL queries to extract, join, and relationalize multi-dimensional customer and sales data from enterprise local databases — replacing slow, fragile ad-hoc scripts.

    Stable extracts at a fraction of the time

  3. 03

    ETL Pipeline Engineering

    Engineered an automated ETL workflow in Python (Pandas) to consolidate disparate monthly CSV files into a single trusted source, eliminating manual processing.

    ≈ 5 hours reclaimed per week

  4. 04

    Business Intelligence & Visualization

    Designed and deployed dynamic tracking dashboards to monitor KPIs, translating raw metrics into actionable insight for stakeholders across the business.

    Built-in BI for decision makers

  5. 05

    Cross-Functional Data Storytelling

    Partnered with team members to communicate complex analytical findings, driving strategic decisions through clear, evidence-first storytelling.

    Findings → decisions, not reports

Projects

04 / Selected work

Four products, each taken the whole way — from a messy first source to a decision someone could act on.

0108/2025 – Current

Flying Sapphire

Agentic AI Web Browser — Co-Founder & AI Workflow Developer

A browser that thinks. Local LLMs route tasks, read active tabs, and execute multi-step work across 40+ enterprise APIs — with no cloud computation.

The problem

Typical AI assistants ship your context to the cloud and stop at chat. Sapphire had to act on what you're actually looking at — live tabs, live data — and do the work end-to-end, privately, on the machine itself.

Approach

  • Core agentic logic

    Integrated local LLMs and engineered the agent's reasoning loop so prompts are parsed into intent and routed to the right tool with no cloud round-trip.

  • Data extraction pipelines

    Built parsing pipelines that distill raw, unstructured web text into clean, structured Markdown study guides and structured JSON.

  • API routing layer

    Wired the agent to 40+ external enterprise APIs — Google Workspace, Slack, logistics tooling — for autonomous logistics and data-entry tasks.

  • Latency & context tuning

    Tuned system prompts and context windows so the agent reads live tabs, debugs Python, and produces insights at near-zero latency.

PythonLLMsAgentic WorkflowsAPI IntegrationPrompt EngineeringMarkdown / JSON Structuring
Compute
On-device, local
APIs wired
40+
Output format
JSON / Markdown
Role
Co-Founder

Outcomes

  • No cloud compute

    All routing and decisioning runs on-device; privacy is the default, not a setting.

  • 40+ integrations live

    A single natural-language trigger can span calendars, docs, and messaging.

— The browser that brings its own hands.

0204/2026 – 07/2026

Sports Odds Predictor

Forecasting & Anomaly Detection Pipeline

An end-to-end ML pipeline that forecasts match outcomes and flags irregular betting patterns — a simulated automated fraud and risk detection stack.

The problem

Bookmaking data is messy, interdependent, and adversarial. The task was to build the whole risk stack from scratch: forecast outcomes, then catch the bet patterns that do not behave like the model expects.

Approach

  • Predictive core

    Built a Python + Scikit-learn pipeline to forecast match outcomes from historical transactional data, engineered for drift rather than best-case accuracy.

  • Feature engineering

    Used Pandas and NumPy to clean raw betting datasets, impute missing values, and engineer rolling statistical features.

  • Risk detection

    Implemented classification models to surface high-risk betting anomalies and flag them for automated review.

  • Storage & retrieval

    Structured and queried T-SQL tables to store and process multi-dimensional match and odds metrics efficiently.

PythonMLflowScikit-learnPandasMLOpsReal-Time AnalyticsT-SQL
Task
Classification
Tracking
MLflow
Data
Streaming simulation
Store
T-SQL

Outcomes

  • End-to-end simulated stack

    Ingestion → features → forecast → risk flag, tracked through MLflow.

  • Streaming-ready

    Designed so a real-time odds feed can be dropped into the pipeline unchanged.

— Predict the game. Police the bet.

0304/2026 – 05/2026

Wine Quality Model

Predictive Modeling & A/B Testing

Rigorous A/B testing paired with a predictive pipeline classifying wine quality metrics. Scored a perfect 100% as a university assignment.

The problem

A deceptively simple dataset hides a demanding question: which chemical properties actually move quality, and which differences are just noise? The brief demanded statistical honesty, not a leaderboard win.

Approach

  • Data engineering

    Extracted, joined, and preprocessed structured datasets from both CSV and JSON sources, keeping provenance intact.

  • Statistical design

    Designed and executed A/B tests to statistically validate hypotheses on chemical properties and their effect on quality scores.

  • Full lifecycle trace

    Documented the complete ML lifecycle — cleaning to evaluation — in Jupyter Notebooks for reproducibility.

PythonJupyterPandasScikit-learnHypothesis Testing
Score
100%
Method
A/B testing
Inputs
CSV + JSON
Notebook
Experiments.ipynb

Outcomes

  • Perfect score

    100% on the assignment — the validation that A/B rigor transfers to production behavior.

  • Reproducible by design

    Every figure in the report can be regenerated from the notebooks.

— Statistics can taste too.

0401/2026 – 03/2026

Credit Card Approval

Classification & Hyperparameter Tuning

A classification pipeline that automates credit approval decisions from historical financial data, tuned to cut false positives.

The problem

Credit data is brutally unbalanced and full of missing fields. A naive model over-approves and bleeds risk; an over-tuned one quietly rejects good customers. The brief: automate the decision without inheriting either failure.

Approach

  • Preprocessing

    Cleaned the highly unbalanced financial dataset by imputing missing values, encoding categoricals, and scaling features for stability.

  • Systematic tuning

    Used GridSearchCV to optimize Logistic Regression parameters, prioritizing precision and reducing false positives.

PythonNumPyScikit-learnLogistic RegressionGridSearchCV
Model
Logistic Regression
Search
GridSearchCV
Balance
Imputation + scaling
Target
Binary approval

Outcomes

  • Precision-led approvals

    Hyperparameters chosen against false positives, not raw accuracy.

  • A reusable classifier

    Pipeline structure that ports directly to other binary approval tasks.

— Fewer bad yeses. No good noes.

Try the model

Interactive — 05 / Playground

Every threshold is a trade. Move it left and you catch every anomaly plus half the honest traffic. Move it right and the quiet fraud slips through. This is the exact judgment a risk model encodes. 

Anomaly detection — live demo

Drag the risk threshold

flagged: 7 / 30
est. precision: 79%
Sensitive · T=20Threshold 60Conservative · T=95

The way I work

Research note — 06 / Method

Paper — internal note

From Raw to Rendered

A personal methodology for data products

Arjun KhannaApplied Data Sciences, MCAST

Abstract

Guided by coursework in applied data sciences and hardened by production work at GRID and on agentic systems, I treat every engagement as a pipeline, not a script. Problem framing precedes acquisition; evaluation precedes deployment; and stakeholder legibility is a first-class output, never an afterthought. This note records the sequence I run — and the failure modes it is designed to prevent.

cite this method

@misc{khanna2026pipeline,
  author       = {Khanna, Arjun},
  title        = {From Raw to Rendered: A Personal Methodology for Data Products},
  year         = {2026},
  school       = {MCAST — Applied Data Sciences},
  note         = {Compiled from GRID production work and Level 6 coursework}
}
  1. 01

    Problem framing

    Decide what decision this data is meant to support before touching a value. A model without a downstream decision is a technical exercise.

  2. 02

    Acquisition & wrangling

    Find every source — CSV, JSON, SQL, API. Join them deliberately, gate data quality, and make provenance reproducible. This is where pipelines are won or lost.

  3. 03

    Analysis & modeling

    Explore to build hypotheses, then model to test them. Choose the simplest model that clears the bar; reserve complexity for where it earns its keep.

  4. 04

    Evaluation & iteration

    Track everything in MLflow. Prefer precision and business-relevant metrics over leaderboard accuracy. Iterate on the metric the stakeholder feels, not the one the notebook loves.

  5. 05

    Communication & automation

    Rendered output — dashboard, brief, or agent — that a stakeholder can act on. Then automate the pipeline so the next refresh does not need a human babysitter.

The toolbox

07 / Capabilities

Languages & APIs01
  • ·Python
  • ·SQL
  • ·T-SQL
  • ·C#
  • ·Excel
  • ·API Integration
  • ·DSA
  • ·DBMS
Libraries & Frameworks02
  • ·Pandas
  • ·NumPy
  • ·Matplotlib
  • ·Scikit-Learn
  • ·PyTorch
  • ·XGBoost
  • ·FastAPI
ML & AI03
  • ·Machine Learning
  • ·Predictive Modeling
  • ·Linear Regression
  • ·LLMs
  • ·Agentic AI
  • ·Prompting
MLOps04
  • ·GridSearchCV
  • ·Feature Scaling
  • ·Data Imputation
  • ·MLOps
  • ·Real-Time Analytics
Statistics05
  • ·A/B Testing
  • ·Hypothesis Testing
  • ·t-Test
  • ·Chi-Square
  • ·Pearson Correlation
  • ·Decision Analysis
BI & Analytics06
  • ·Power BI
  • ·Descriptive
  • ·Diagnostic
  • ·Prescriptive
  • ·Data Visualization
  • ·Revenue & Cost Analysis
Data Management07
  • ·Data Governance & Security
  • ·JSON / Markdown Structuring
  • ·Git / Version Control
Professional08
  • ·Analytical & Critical Thinking
  • ·Problem Solving
  • ·Leadership
  • ·Data Storytelling
  • ·Teamwork

01

Languages & APIs

  • Python
  • SQL
  • T-SQL
  • C#
  • Excel
  • API Integration
  • DSA
  • DBMS

02

Libraries & Frameworks

  • Pandas
  • NumPy
  • Matplotlib
  • Scikit-Learn
  • PyTorch
  • XGBoost
  • FastAPI

03

ML & AI

  • Machine Learning
  • Predictive Modeling
  • Linear Regression
  • LLMs
  • Agentic AI
  • Prompting

04

MLOps

  • GridSearchCV
  • Feature Scaling
  • Data Imputation
  • MLOps
  • Real-Time Analytics

05

Statistics

  • A/B Testing
  • Hypothesis Testing
  • t-Test
  • Chi-Square
  • Pearson Correlation
  • Decision Analysis

06

BI & Analytics

  • Power BI
  • Descriptive
  • Diagnostic
  • Prescriptive
  • Data Visualization
  • Revenue & Cost Analysis

07

Data Management

  • Data Governance & Security
  • JSON / Markdown Structuring
  • Git / Version Control

08

Professional

  • Analytical & Critical Thinking
  • Problem Solving
  • Leadership
  • Data Storytelling
  • Teamwork
Agentic AIC#Power BIMLflowFastAPIXGBoostPyTorchNumPyPandasScikit-LearnSQL / T-SQLPythonAgentic AIC#Power BIMLflowFastAPIXGBoostPyTorchNumPyPandasScikit-LearnSQL / T-SQLPython

Education & training

BSc (Honours) in Applied Data Sciences

Applied statistics, machine learning, data engineering, and decision science — studied in the lab and pressure-tested on real pipelines.

Institution
MCAST
Level
MQF / EQF Level 6
Period
01/10/2025 – Current
Location
Iż-Żurrieq, Malta
mcast.edu.mt ↗

6

MQF/EQF level 6 (honours)

100%

A/B testing ML assignment

4+

Models taken to production

Language skills

LanguageListenReadSpeakInteractWrite
HindiMother tongue—————
EnglishProficient userC2C2C2C2C2

Levels: A1/A2 basic user · B1/B2 independent user · C1/C2 proficient user.

08 / Contact

Let's buildsomething that predicts.

A pipeline, a prediction, a dashboard — or a browser that does the work for you. If the data exists, there is a way to make it decide. Let's find it.