Python Sets: How I Learned to Work with Unique Values
When I first came across sets in Python, I didn't immediately see why I needed them. I already had lists and dictionaries, so I wondered what sets were supposed to do differently. The thing that made sets click for me
When I first came across sets in Python, I didn't immediately see why I needed them.
I already had lists and dictionaries, so I wondered what sets were supposed to do differently.
The thing that made sets click for me was one word , unique.
A set is useful when I want to store values without duplicates.
Creating a set
A set is created using curly brackets.
For example:
numbers = {1, 2, 3, 4}
I can also create a set from a list:
numbers = set([1, 2, 2, 3, 3, 4])
The result is:
{1, 2, 3, 4}
The duplicate values are removed automatically.
That was the first thing that made sets useful to me.
Why would I use a set?
Imagine I have customer locations:
locations = ["Nairobi", "Mombasa", "Nairobi", "Kisumu", "Mombasa"]
If I want to know the unique locations, I can do:
unique_locations = set(locations)
Now I get:
{"Nairobi", "Mombasa", "Kisumu"}
Instead of manually checking for duplicates, Python handles it for me.
Adding items
I can add an item using add().
locations = {"Nairobi", "Mombasa"}
locations.add("Kisumu")
Now the set contains:
{"Nairobi", "Mombasa", "Kisumu"}
If I try to add something that's already there:
locations.add("Nairobi")
Python doesn't create another "Nairobi".
That's because sets only keep unique values.
Removing items
I can remove an item using remove():
locations.remove("Mombasa")
I can also use discard():
locations.discard("Mombasa")
The difference is that remove() raises an error if the item doesn't exist, while discard() simply does nothing.
I found discard() useful when I wasn't sure whether a value was already in the set.
Checking if an item exists
Just like with lists, I can use in.
For example:
locations = {"Nairobi", "Mombasa", "Kisumu"}
print("Nairobi" in locations)
The result is:
True
This is useful when I only need to know whether a value exists.
Finding the length
I can use len() with a set:
locations = {"Nairobi", "Mombasa", "Kisumu"}
print(len(locations))
The result is:
3
This tells me how many unique values are in the set.
Sets and loops
I can loop through a set just like other Python collections.
For example:
locations = {"Nairobi", "Mombasa", "Kisumu"}
for location in locations:
print(location)
Python goes through each unique value.
One thing I had to understand is that sets don't work like lists when it comes to positions.
Sets don't use indexing
With a list, I can do:
products = ["Bread", "Milk", "Sugar"]
print(products[0])
But I can't do that with a set:
products = {"Bread", "Milk", "Sugar"}
print(products[0])
A set doesn't have positions that I can rely on.
That is another reason I wouldn't use a set when the order of items matters.
Union
Sets become really useful when comparing groups of data.
Suppose I have customers from two branches:
nairobi_customers = {"Amina", "Brian", "Feddy"}
mombasa_customers = {"Brian", "Faith", "Kevin"}
I can combine all unique customers using union():
all_customers = nairobi_customers.union(mombasa_customers)
The result contains:
{"Amina", "Brian", "Feddy", "Faith", "Kevin"}
Brian only appears once.
Intersection
What if I want to find customers who appear in both branches?
I can use intersection():
common_customers = nairobi_customers.intersection(mombasa_customers)
The result is:
{"Brian"}
This was one of the most useful set operations for me because it makes comparing groups very easy.
Difference
I can also find values that exist in one set but not another.
For example:
nairobi_only = nairobi_customers.difference(mombasa_customers)
The result is:
{"Amina", "Feddy"}
These are customers in the Nairobi set who are not in the Mombasa set.
Symmetric difference
There is also symmetric_difference().
It gives me values that are in either set, but not in both.
different_customers = nairobi_customers.symmetric_difference(mombasa_customers)
The result is:
{"Amina", "Feddy", "Faith", "Kevin"}
Brian isn't included because he appears in both sets.
A practical example
Imagine I have two lists of students who attended different events.
event_one = ["Feddy", "Amina", "Brian", "Faith"]
event_two = ["Brian", "Faith", "Kevin", "David"]
I can convert them to sets:
set_one = set(event_one)
set_two = set(event_two)
Now I can find students who attended both events:
set_one.intersection(set_two)
Students who only attended the first event:
set_one.difference(set_two)
And everyone who attended at least one of the events:
set_one.union(set_two)
This is where I started seeing why sets exist.
They're not just another way to store a bunch of values. They're really useful when I care about uniqueness and comparing groups.
Converting a set back to a list
Sometimes I may need a list after removing duplicates.
For example:
names = ["Feddy", "Amina", "Feddy", "Brian"]
I can remove duplicates:
unique_names = set(names)
And convert the result back to a list:
unique_names = list(set(names))
Now I have a list containing only the unique names.
What I understood
The main thing I learnt about sets is that I shouldn't think of them as a replacement for lists.
They solve a different problem.
If I need ordered data, I think of a list.
If I need unique data, I think of a set.
If I need to compare two groups of values, sets become especially useful.
Final thoughts
Sets were probably one of those Python topics that only made sense after I actually had a problem to solve.
The moment I needed to remove duplicate data or compare two groups, sets suddenly became useful.
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes ā full credit and traffic to the original publisher.