Technology

Data Scientist Interview Questions and Answers

Data science and machine learning interviews mix statistics, ML concepts, practical data handling and, increasingly, generative AI. Expect to explain models in plain language and to justify your metrics and choices. These questions cover the fundamentals through to model deployment and RAG.

Reading answers is not the same as saying them.Practise data science & ai questions out loud and get a score, what you missed and a model answer for each one.
Practise free with AI

Topics interviewers ask about

PythonPandasNumPyStatisticsMachine LearningDeep LearningTensorFlowPyTorchscikit-learnNLPComputer VisionGenerative AI & LLMsData VisualizationPower BITableauExcelSQL for Analytics

Basic data science & ai interview questions

Fundamentals, definitions and simple scenarios. Good for freshers and warm-ups.

1. What is the difference between supervised and unsupervised learning?

Supervised learning trains on labelled data to predict a known target: classification predicts a category (spam or not) and regression predicts a number (house price). Unsupervised learning has no labels and finds structure in the data, such as clustering customers into segments or reducing dimensions with PCA. Reinforcement learning, a third type, learns by trial and error from rewards.

2. When would you use the median instead of the mean?

The mean is pulled by outliers and skewed data, while the median, the middle value, is not. For incomes, house prices or delivery times, a few extreme values can make the mean misleading, so the median better describes a typical value. For roughly symmetric data without outliers, the mean is fine and uses all the information.

3. What is overfitting, and how do you prevent it?

Overfitting is when a model learns noise in the training data, so it scores well on training data but poorly on new data. I detect it by comparing training and validation scores. To prevent it I use more data, a simpler model, regularisation (L1 or L2), cross-validation, early stopping, dropout in neural networks, and pruning or depth limits for trees.

4. Explain precision and recall.

Precision is TP / (TP + FP): of everything the model flagged as positive, how much was right. Recall is TP / (TP + FN): of all real positives, how many the model found. Changing the decision threshold trades one for the other, and F1 is their harmonic mean. For disease screening I favour recall, so cases are not missed; for spam filtering, precision, so real emails are not blocked.

Intermediate data science & ai interview questions

Applied problems, trade-offs and questions about your own projects.

5. Explain the bias-variance trade-off.

Bias is error from a model that is too simple to capture the pattern, which causes underfitting. Variance is error from a model so sensitive to its training data that it changes a lot with small data changes, which causes overfitting. Making a model more complex usually lowers bias and raises variance. The goal is the lowest total error on unseen data, which I find with validation curves and cross-validation.

6. How do you handle missing data?

First I ask why it is missing, because data missing at random is handled differently from data missing for a reason. If very few rows are affected I may drop them. Otherwise I impute with the median or mode, or model-based methods like KNN or iterative imputation, and often add a flag column showing the value was missing. The imputer must be fitted only on training data to avoid leakage.

7. How do you deal with an imbalanced dataset?

Accuracy is misleading when, for example, 99% of transactions are not fraud, so I use precision, recall, F1 and PR-AUC. Then I try class weights in the model, resampling of the training set only (undersampling, oversampling or SMOTE), and tuning the decision threshold for the business cost of errors. Collecting more minority-class examples is often the best fix.

8. What is the difference between Random Forest and Gradient Boosting?

A Random Forest trains many deep trees independently on bootstrap samples with random feature subsets and averages them, which mainly reduces variance; it is robust and needs little tuning. Gradient Boosting (XGBoost, LightGBM, CatBoost) builds trees one after another, each correcting the previous ones' errors, which reduces bias. It often gives better accuracy on tabular data but needs careful tuning and can overfit.

High level data science & ai interview questions

System design, deep internals, leadership and tough follow-ups.

9. How does the attention mechanism in transformers work?

Each token is turned into a query, a key and a value vector. Attention scores are softmax(QKᵀ / √d), and each token's output is a weighted sum of the value vectors, so every token can draw on context from every other token, in parallel. Multi-head attention learns several kinds of relationships at once, positional encodings add word order, and layers stack attention with feed-forward networks. GPT-style models use decoder-only transformers; BERT uses the encoder.

10. What is data leakage? Give examples.

Leakage is when information that will not be available at prediction time gets into training, so validation looks great but the model fails in production. Examples are scaling or imputing on the full dataset before splitting, features derived from the target, duplicate records across train and test, and using future data in time series. I prevent it with pipelines fitted inside each training fold and time-based splits.

11. How would you deploy and monitor a machine learning model?

I package the preprocessing and model together as one pipeline, version it in a tool like MLflow, and serve it through an API or batch jobs. In production I monitor input data drift, prediction drift, latency and errors, and real performance once labels arrive. A new model is compared with the current one through shadow or A/B testing, and drift or a performance drop triggers retraining.

12. What is RAG (retrieval-augmented generation)?

RAG gives a large language model access to your own data. Documents are split into chunks, converted to embeddings and stored in a vector database. For each question the most relevant chunks are retrieved, often reranked, and added to the prompt, so the model answers from that context and can cite sources. Quality depends on chunking, the embedding model and retrieval, and it must be evaluated for groundedness to catch hallucinations.

Ready to test yourself?Pick your topics and level, answer by voice or text, and get instant feedback. Free.
Start a mock interview

More technology interview questions