How do I understand about the landscape of Data Science Techniques?
This article is simply an attempt to organize everything I've been learning in one place. When I first started learning data science, I was overwhelmed by the number of techniques, algorithms, and buzzwords. Every new t
This article is simply an attempt to organize everything I've been learning in one place.
When I first started learning data science, I was overwhelmed by the number of techniques, algorithms, and buzzwords. Every new topic seemed to introduce another family of methods!!
The list goes like beloww
regression, classification, clustering, dimensionality reduction, recommendation systems, neural networks... the list just kept growing.
I'm writing this from the perspective of someone who has only recently stepped into this field. Yes, we're already living in the AI era, and sometimes it feels like I'm very behind. But I've realized that one of the best ways to learn is to document the journey.
So, I started asking myself a few simple questions:
What exactly are we trying to solve?
Why are there so many techniques in data science?
Why do we need dozens of algorithms for seemingly similar problems?
How do data scientists decide which technique to use?
The answer begins with understanding the problem, not the algorithm.
Many predictive machine learning tasks can be grouped into two broad categories:
Classification – predicting a category or class (for example, whether an email is spam or not).
Regression (Function Approximation) – predicting a continuous numerical value (for example, forecasting house prices).
However, data science is much broader than prediction alone. We also encounter problems such as clustering similar customers, detecting fraudulent transactions, reducing the dimensionality of data, forecasting future trends, recommending products, processing text, analyzing images, and much more.
Before choosing any technique, we first need to understand the data itself.
In the real world, data is rarely clean or perfectly structured. Businesses generate massive amounts of information every second!
From application logs and server metrics to customer transactions, clickstreams, sensor readings, and social media activity. In many cases, this isn't just "data"; it's Big Data.
Some datasets are structured and easy to work with, while others are noisy, incomplete, inconsistent, or completely unstructured.
That's why the first job of a data scientist isn't to build a machine learning model but it's to understand the data and the business problem. Once both are clear, we can identify the type of problem we're solving and select the most appropriate family of techniques.
In this article, I'll explore these different categories of problems and the techniques commonly used to solve them.
Classification Problems
Here, We deal with labled data, where the goal is to assign a label or class to a new data point based on learned patterns. Classfication also further have several types.
Binary Classification-
Has only 2 label. you have data that is already labeled. Our Model is Trained on it. Now, New data point will come and the new data point will a unlabled data. The class of that data will be pridicted based on the features it has and will be given a type.
Example - Fraud Detection
Two Classes
Every transaction has various features such as
New Data comes in X
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.


