GATE DA syllabus 2027
Data Science & Artificial Intelligence
GATE's newest and fastest-growing paper, blending statistics, core CS and machine learning.
Syllabus sections
- Probability & Statistics
- Linear Algebra
- Calculus & Optimization
- Programming, Data Structures & Algorithms
- Database Management & Warehousing
- Machine Learning
- Artificial Intelligence
Subject-wise weightage (indicative)
| Topic | Typical marks |
|---|---|
| General Aptitude | 15 |
| Probability & Statistics | 15+ |
| Machine Learning | 12–15 |
| Linear Algebra | 7–9 |
| Programming, DS & Algorithms | 7–9 |
| Artificial Intelligence | 6–8 |
| Database Management & Warehousing | 5–7 |
| Calculus & Optimization | 5–7 |
Weightage is indicative, based on recent papers. It varies year to year.
About the GATE DA paper
Data Science & Artificial Intelligence was introduced in 2024, which makes it the newest paper in GATE and the one with the least previous year material available. That single fact shapes how it should be prepared. In an older paper you can lean on twenty years of PYQs to infer the examiner’s taste. In DA you cannot, so you have to prepare the syllabus as written rather than the syllabus as historically asked.
The paper’s centre of gravity is mathematical, not systems oriented. Probability, statistics and linear algebra are not a preliminary section here the way Engineering Mathematics is in other papers. They are the paper. Add Machine Learning and those four units account for the clear majority of the subject marks.
One structural point that catches candidates out: DA has no Engineering Mathematics section at all. Its mathematics is distributed into three named units with their own, deeper syllabi.
What each section of the syllabus actually covers
Probability & Statistics
The largest unit in the paper. It covers counting through permutations and combinations, the probability axioms, sample spaces and events, independent and mutually exclusive events, and marginal, conditional and joint probability. Bayes’ theorem appears repeatedly, as do conditional expectation and variance.
Descriptive statistics covers mean, median, mode, standard deviation, correlation and covariance. Random variables are treated in depth: discrete variables with uniform, Bernoulli and binomial mass functions, and continuous variables with uniform, exponential, Poisson, normal and standard normal distributions, plus the t and chi-squared distributions. Cumulative distribution functions and conditional density functions are examinable.
The inference tail of this unit is what most candidates underestimate: the central limit theorem, confidence intervals, and the z-test, t-test and chi-squared test. These are mechanically simple once learned, and they are asked.
Linear Algebra
Vector spaces and subspaces, linear dependence and independence, matrices, and the special matrix families the paper names explicitly: projection, orthogonal, idempotent and partitioned matrices, each with their properties. Quadratic forms are included.
On the computational side: systems of linear equations and their solutions, Gaussian elimination, determinants, rank and nullity, eigenvalues and eigenvectors, projections, LU decomposition and singular value decomposition.
SVD and projections deserve emphasis because they connect directly to the machine learning unit, where principal component analysis is essentially this material wearing a different name.
Calculus & Optimization
The most compact unit in the paper: functions of a single variable, limits, continuity and differentiability, Taylor series, maxima and minima, and optimisation involving a single variable.
Note the scope carefully. This unit stays in one variable. Its real purpose is to support the optimisation intuition behind regression and gradient based learning, so it is worth studying with that connection in mind rather than as isolated calculus.
Programming, Data Structures & Algorithms
Programming here is Python, not C, which is a meaningful difference from the CS paper. The data structures are the standard set: stacks, queues, linked lists, trees and hash tables. Search covers linear and binary search. Sorting covers the elementary algorithms of selection, bubble and insertion sort, plus the divide and conquer pair of mergesort and quicksort. Graphs are introduced along with traversals and shortest path algorithms.
The depth expected is noticeably lower than in GATE CS. There is no compiler design, no advanced dynamic programming, no amortised analysis. If you are preparing DA alongside CS, this unit is nearly free.
Database Management & Warehousing
The database half is familiar: the ER model, the relational model with relational algebra, tuple calculus and SQL, integrity constraints, normal forms, file organisation and indexing.
The warehousing half is the part with no CS equivalent, and it is where marks are quietly lost. It covers data types and data transformation including normalisation, discretisation, sampling and compression, then data warehouse modelling: schemas for multidimensional data models, concept hierarchies, and measures with their categorisation and computation. Because this material does not appear in any other GATE paper, there is no shortcut through it.
Machine Learning
The second largest unit, split into supervised and unsupervised learning.
Supervised learning covers regression and classification: simple and multiple linear regression, ridge regression, logistic regression, k-nearest neighbour, the naive Bayes classifier, linear discriminant analysis, support vector machines, decision trees, the bias-variance trade-off, cross validation methods including leave-one-out and k-fold, the multi-layer perceptron and feed-forward neural networks.
Unsupervised learning covers clustering with k-means and k-medoid, hierarchical clustering in both top-down and bottom-up form with single and multiple linkage, and dimensionality reduction through principal component analysis.
Questions here are rarely conceptual essays. They tend to be small numerical exercises: run one step of k-means by hand, compute a decision tree split, work out a bias-variance decomposition, or evaluate a classifier on a small confusion matrix.
Artificial Intelligence
Three strands. Search covers uninformed, informed and adaptive search. Logic covers propositional and predicate logic. Reasoning under uncertainty covers conditional independence representation, exact inference by variable elimination, and approximate inference by sampling.
The uncertainty strand is the one that connects back to the probability unit, and Bayesian network questions asking for a conditional independence judgement or a variable elimination ordering are a natural fit for this paper.
Reading the weightage table
DA concentrates marks more tightly than most GATE papers. Probability & Statistics and Machine Learning together are typically worth close to thirty marks. Add General Aptitude at a fixed 15 and roughly forty-five of the hundred marks sit in three areas.
That concentration cuts both ways. It means a candidate with a strong statistics background starts a long way ahead. It also means there is nowhere to hide: you cannot skip probability in DA the way a CS candidate might skip compiler design and still be competitive.
Because the paper is new, treat any published weightage, including the table above, as indicative rather than settled. Three cycles is not enough history to call a pattern.
A preparation order that works
- Probability and statistics first, and slowly. Everything downstream leans on it, including naive Bayes, linear discriminant analysis, the bias-variance decomposition and the whole uncertainty strand of AI.
- Linear algebra next, with projections, eigen-decomposition and SVD done properly rather than skimmed. Principal component analysis becomes almost free afterwards.
- Calculus and optimisation, kept short and tied to the regression material.
- Machine learning, once the three mathematical units are solid. Learning it before the maths produces recall without the ability to compute, which is the wrong side of the trade for this exam.
- Programming, data structures and algorithms, which is self-contained and can be slotted in whenever convenient.
- Databases and warehousing, leaving deliberate time for the warehousing topics that have no equivalent elsewhere.
- Artificial intelligence last, since its uncertainty section builds on probability.
- General Aptitude throughout, in short weekly sessions.
Preparing a paper with only three years of PYQs
With papers from 2024 onwards only, the usual strategy of drilling twenty years of previous questions is unavailable. Three adjustments follow.
First, solve every available DA paper carefully rather than quickly. A small corpus should be studied, not consumed.
Second, borrow from adjacent papers where the syllabus genuinely overlaps. GATE CS previous year questions cover the databases, data structures and algorithms material well. GATE ST questions cover probability, distributions and inference at a comparable level. This is supplementary practice, not a substitute, and it needs filtering against the DA syllabus because both papers go deeper in places DA does not.
Third, lean harder on full-length mock tests than a CS or ME candidate would. In an established paper, mocks calibrate speed. In DA they are also your main source of unseen questions.
Where candidates lose marks
- Treating statistical inference as optional. Confidence intervals and hypothesis tests are explicitly in the syllabus and are the most commonly skipped topic in the paper.
- Studying machine learning as vocabulary. Knowing what ridge regression is does not help if you cannot compute with it on a four-point dataset.
- Ignoring data warehousing. It looks peripheral, it is unfamiliar, and it appears nowhere else in GATE, which is exactly why it is worth the hours.
- Assuming the CS syllabus transfers wholesale. The overlap is real but partial, and the programming language here is Python.
- Negative marking on guesses. MCQs carry one-third and two-thirds penalties. Numerical Answer Type questions carry none, so an NAT question left blank is a strictly worse choice than an attempt.
General Aptitude, the section nobody should concede
General Aptitude is 15 marks in every GATE paper, needs no subject background, and covers verbal ability, numerical reasoning, data interpretation and basic quantitative aptitude. In a paper as mathematically concentrated as DA it is tempting to treat it as trivially easy and leave it to exam day. Book a short weekly slot for it instead. Fifteen marks banked reliably outweighs a fourth pass through the machine learning notes.
Recommended books for GATE DA
- Pattern Recognition and Machine Learning by Christopher Bishop
- An Introduction to Statistical Learning by James, Witten, Hastie, Tibshirani
- Artificial Intelligence: A Modern Approach by Russell & Norvig
Frequently asked questions
When was the GATE DA paper introduced?
Data Science & Artificial Intelligence (DA) was introduced in 2024, so only a few previous year papers exist. Mock tests and topic-wise practice matter more in DA than in older papers.
Can I appear for both GATE DA and GATE CS?
Yes. GATE allows up to two papers from an approved combination list, and DA with CS is the most popular pairing because of heavy syllabus overlap in programming, data structures, algorithms and databases.
Does GATE DA have a separate Engineering Mathematics section?
No. DA does not have the standard Engineering Mathematics section. Instead it has its own Probability & Statistics, Linear Algebra, and Calculus & Optimization units, which together carry very high weightage.
Is GATE DA easier than GATE CS?
Neither is reliably easier. DA leans heavily on probability, statistics and machine learning intuition, while CS is broader across systems subjects. Choose based on your background rather than perceived difficulty.