Introduction to Network-Based Exploit Development
The first step here is to develop a C server, a tiny one of course, with safety features turned off, and start it on localhost. But we need to learn one thing first: what exactly is a buffer overflow. Buffer Ov
The first step here is to develop a C server, a tiny one of course, with safety features turned off, and start it on localhost. But we need to learn one thing first: what exactly is a buffer overflow.
Buffer Overflow
A buffer overflow happens when a program writes more data to a fixed-size memory container, which is called a buffer. So anything more than what it is supposed to hold, and boom: overflow. Memory is contiguous, so guess what, the extra data does not go away. It spills over and overwrites whatever data is next to it in RAM. It's like trying to pour 5 gallons of water into a 1-gallon bucket. It floods the floor around it.
The Stack Memory Diagram
When a function runs, it sets aside a "stack frame" in memory, which is a temporary chunk of memory that a computer sets aside every single time a function is called.
1. Normal State (Everything fits)
+------------------------+
| Saved Return Address | -> Tells the CPU where to return when the function finishes
+------------------------+
| Local Buffer (64 bytes)| -> The container meant for your input data
+------------------------+
---
2. Overflowed State (Too much data sent)
+------------------------+
| [ A A A A A A A A ] | -> OVERWRITTEN! The return address is now corrupted with "AAAA" (0x41414141)
+------------------------+
| [ A A A A A A A A ] | -> Spilled over past the 64-byte boundary
+------------------------+
| Local Buffer (64 bytes)| -> Completely full of data
+------------------------+
So why does it cause a crash?
When the function is done running, the CPU looks at the saved return address to figure out what code to execute next. Now, because we have replaced that data with a bunch of A's, it looks at that memory address instead. Where is that memory address coming from? Well, I wrote it, right? When I sent a flood of data over the network, like a hundred A's, the input filled up the 64-byte buffer and kept spilling over and over. So whatever the CPU went back to do at the saved return address is now gone: the extra A's took it out and erased it. A equals 41 in hex. In computer language, the capital letter A is represented by the hexadecimal value 41. So when you sent four A's, you literally wrote 41 41 41 41 (0x41414141) straight into that memory slot. So the CPU now goes to that memory address and then crashes. A is also 65 in decimal, according to the ASCII table, that is not a magical number we can find but it is a a table we look up. So, when you convert 65 to hex it becomes 41. The math is quite simple.
Math:
Divide 65 by 16 and you will get = 4 remainder 1.
The 4 is the first hex digit. 1. The 1 is the second hex digit. 65 decimal = 41 hex
65 Γ· 16 = 4 remainder 1
β β
4 1
A = 41 in hex
So when we do simple division we get 65 Γ· 16 = 4.0625. But for hex, we do whole-number division:
16 fits inside 65 4 times
16 Γ 4 = 64
65 - 64 = 1
So:
65 Γ· 16
16 Γ 4 = 64
65 - 64 = 1
= 4 with 1 left over
4 and 1 β 41 Hex.
Summary:
65 Γ· 16
16 goes into 65
4 times
16 Γ 4 = 64
65 - 64 = 1
FIRST DIGIT = 4
SECOND DIGIT = 1
β
HEX = 41
ASCII says:
65 = A
So:
A = 65 decimal = 41 hex
We do need to understand hex (hexadecimal) to the tee, because in exploit development, hex is the native language. Every memory address, register value, and bad character you look at will be in hex. It kills my brain cells sometimes to figure out hex, but before we get into hex we need to understand what decimal is, or BASE-10.
Decimal (Base-10) and Hex (Base-16)
Everyday counting we do from 0 to 9, but in digital form basically. 0-9 is 10 numbers, because 0 is the first number in computer language. Hex, on the other hand, uses 16 symbols. Decimal runs out of symbols after 9, so that is where hex comes in: after 9 comes A, B, C, D, E, F. Then it rolls over.
We have to memorize these numbers: A=10, B=11, C=12, D=13, E=14, F=15. From here, everything else is just multiplication.
Now let's take 2 digits: 2 and F.
The left digit is the 16s slot, and the right side, F, is the 1s slot. Same as 47 being four 10s and seven 1s.
Steps:
- Left digit is 2. Multiply by 16, so 2 * 16 = 32.
- Right digit is F. F is just 15 (from the memorized list). It's the 1s slot, so 15 * 1 = 15.
- 32 + 15 = 47. So 2F is 47.
That's the whole trick every single time: left digit times 16, plus right digit.
So let's decipher what 3A is:
3 * 16 = 48 and A * 1 = 10 * 1 = 10. So 48 + 10 = 58.
Now, maybe another day we can do more math. For now, let's actually break some things, but that's not gonna happen yet, right? We need to know what a BIT is.
A bit is one switch, which means it's either 0 or 1. Just yes or no, and that is it. So we line up 8 of them in a row. Now that just became a byte. Each switch in the row is worth double the one to its right, starting at 1 on the far right, and at the end count it up, it will be 255.
Now how do we do the byte math? To read a byte, whenever there is a 1, you grab that slot's value. Then add them up. You skip the slot with 0. Bytes max out at 255.
SIMPLE RULE -> Write the 8 slot values. Look at the bits. Keep the slots with 1 and skip the slots with 0, boom. Let's do simple mathematics.
Take the byte 01000001.
Line the bits under their slots:
128 64 32 16 8 4 2 1
0 1 0 0 0 0 0 1
Slots with 1 can be kept only, anything else is gone, so now I can see we can only keep 1 and 64, which becomes 64+1 = 65. So 01000001 is 65 and 65 is also the letter A.
65 is decimal, not hex. The byte 01000001 adds up to 65 in decimal. Same byte in hex would be 41. Almost like it's Halloween. 3 costumes.
Now the part you're really asking: why does that value mean the letter A?
Because somebody decided it would. There's a lookup table called ASCII that maps numbers to characters. It's just an agreed-upon list. On that list, the number 65 is assigned to capital A. 66 is B, 67 is C, and so on. No math makes 65 "become" A. It's a dictionary, and the computer looks it up.
So how does 01000001 turn into 41 in hex?
There is a trick here: we divide the 8 into 2, then it becomes 4, and each 4 is a hex. 01000001 splits into 0100 and 0001.
Now read each group of 4 using slot values, but a group of 4 only has these slots: 8, 4, 2, 1.
Left group 0100:
8 4 2 1
0 1 0 0
Only the 4 slot is on. So this group = 4.
Right group 0001:
8 4 2 1
0 0 0 1
Only the 1 slot is on. So this group = 1.
Now add them both up and you get 41. Killed all my brain cells for today, man. That's why hex and bytes are best friends: 4 bits always make exactly one hex digit (0 to F), so 8 bits make exactly two hex digits. A byte is always two hex digits, every time.
The Vulnerable Server
OS: Kali. I saved it as server.c.
#include <stdio.h> // Stands for input and output - printf()
#include <string.h> // String and memory functions
#include <unistd.h> // Gives you the raw system calls like -> read(), write(), and close()
#include <arpa/inet.h> // Gives you network address tools -> htons(), which flips your port number into the byte order the network expects.
int main(){
// Socket is basically the network channel
int sock = socket(AF_INET, SOCK_STREAM, 0), client; // 0 is for accepting the default transport protocol, which is TCP.
// htons just takes the port and converts it into a byte order that the network understands.
// sockaddress is the address for the network channel
struct sockaddr_in Sockaddress = {AF_INET, htons(31337), {0}}; // 0 is open, meaning listen on all my machine's IPs, not just one.
// Bind = attach the address to the network channel.
// sock is of course the socket network channel and we are binding it with the socket address.
// (struct sockaddr*) is a generic label bind accepts; the * makes it -> pointer to.
// &Sockaddress is the actual pointer to the sockaddr_in box we made earlier.
bind(sock, (struct sockaddr*)&Sockaddress, sizeof(Sockaddress));
listen(sock, 1); // Start accepting incoming connections on the socket & 1 = how many callers can wait on hold while you're busy with one.
while((client = accept(sock, 0, 0))){
// Remember we are the caller!
// client is whoever connects to your network channel. The first 0 is the caller's info or address, so we will not log that, and since the caller is us and we are on localhost there is no point saving that. The second 0 -> do not save the size of the IP address either.
// A char is one byte and we created a box holding 64 bytes.
// Now we can't really put 512 bytes in the box; it will overflow.
char buf[64];
read(client, buf, 512); // VULNERABILITY: Reads 512 bytes into a 64-byte buffer
close(client);
}
}
Compile code -> gcc server.c -o server -fno-stack-protector -z execstack -no-pie
Before we continue, read what I have to say first below!
Stack Memory
Stack memory is a region in RAM set aside for running functions. Every time a function is called, it grabs a small slice from there, which is called a FRAME. And why does it do that? Well, to hold its local variables and its return address. And what is the return address? The return address is the CPU's bookmark: when this function finishes, jump back here and keep going, but back to the FRAME. When a function is called, the slice or Frame is freed the instant the function finishes. So basically while the function runs it grabs that FRAME, uses it, then releases it. It is called a stack because frames pile on top of each other and come off in reverse order. LAST ON, FIRST OFF. The stack is scratch memory for functions. Function runs, it gets a slice (a frame) holding its local variables and its return address. Function ends, slice gone.
Example: a notebook. You write top to bottom. One task per line.
- Line 1: Doing taxes. When done, go back to the line where I left off.
- You hit a step that needs a sub-task, so you go to a new line.
- Line 2: Doing the math part of the taxes. When done, go back to line 1.
Each line is a function and also has its own work to do, plus a note that says which line to return to when done. When line 2 was finished, we went back to line 1. So what is the bug in our code? The note and the line sit side by side, so when the work in a line runs long and flows past it, it goes over to the note section and writes over the return note. Now the note points to the wrong line, and you jump somewhere you should not, and you crash, boom!
So now listen, there is a safeguard there called a stack canary. A tripwire between your buffer and the return address that catches an overflow. If it's there, our attack won't work, so we remove it using -fno-stack-protector.
-f -> gcc compiler features
-no -> means no! So it's saying no stack protector. So our attack can take place.
Also -z execstack -> there are 2 things you can do in memory. Store stuff in it (data), and run stuff in it as instructions, like code. Normally the stack is allowed to store but not run code. It's locked, data only. The lock itself is the defense. -z execstack unlocks that safety lock. Now the stack can run code too. So why do we care about this? Well, we can sneak our code into buf, which lives on the stack memory, and then we can make the CPU jump onto it. If the stack is data-only, the CPU refuses to run code on it.
-no-pie -> Position Independent Executable. Okay, now normally the OS loads our app in a random location in virtual memory every time it runs, which makes it much harder for attackers to exploit a buffer overflow or a predictable memory address. But when we put the flag -no-pie, the application loads from a static address. Random position = hard to aim at. Fixed position = you can aim. That's why you turn it off for the lab.
ββ$ gcc server.c -o server -fno-stack-protector -z execstack -no-pie
server.c: In function βmainβ:
server.c:25:9: warning: βreadβ writing 512 bytes into a region of size 64 overflows the destination [-Wstringop-overflow=]
25 | read(client, buf, 512); // VULNERABILITY: Reads 512 bytes into a 64-byte buffer
| ^~~~~~~~~~~~~~~~~~~~~~
server.c:24:14: note: destination object βbufβ of size 64
24 | char buf[64];
| ^~~
In file included from server.c:3:
/usr/include/unistd.h:371:16: note: in a call to function βreadβ declared with attribute βaccess (write_only, 2, 3)β
371 | extern ssize_t read (int __fd, void *__buf, size_t __nbytes) __wur
Gave us a big warning.
After that, when I run ./server it's just sitting there waiting. It is waiting on a caller to reach out on port 31337. We can actually confirm it's listening.
ss -tlnp | grep 31337
βββ(kaliγΏkali)-[~/Desktop/Test]
ββ$ ss -tlnp | grep 31337
LISTEN 0 1 0.0.0.0:31337 0.0.0.0:* users:(("server",pid=647479,fd=3))
Now comes the fun part: the FUZZER
Scenario: connect to our server and send way more than 64 bytes to crash it, proving the overflow reaches the return address.
import socket
# Let's create a TCP / IP socket.
server = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
# Now we connect to the server itself.
server.connect(("127.0.0.1", 31337))
# Now let's send some data. It must be bytes. 513 bytes to be exact!
server.sendall(b"A" * 513) # Boom, because A is exactly one byte.
# Clean up and close the connection.
server.close()
Run it -> python3 fuzz.py
Here's the catch with fuzz.py: when you run it, it won't actually crash, because the overflow smashed main()'s return address, but that only matters when main() ends. main() never ends, so it loops forever to accept connections. So the smashed address just sits there unused. No use, no crash. And we do not want to end main(), it needs to keep going to accept connections.
Let's Fix Our server.c Now
FIX: put the vulnerable buffer in a small helper function that ends on every connection it makes, so main() can keep running.
PLAN: make a handle function holding buf and read. main()'s loop calls it, the handle function finishes after each connection, and only then does it read its smashed return note and crash.
#include <stdio.h> // Stands for input and output - printf()
#include <string.h> // String and memory functions
#include <unistd.h> // Gives you the raw system calls like -> read(), write(), and close()
#include <arpa/inet.h> // Gives you network address tools -> htons(), which flips your port number into the byte order the network expects.
// This function handles the connection and holds the vulnerable buffer overflow logic.
void handle_connection(int client_sock){
// void -> this function hands nothing back when it finishes.
// client_sock -> the client's network channel.
// A char is one byte and we created a box holding 64 bytes.
// Now we can't really put 512 bytes in the box; it will overflow.
char buf[64];
read(client_sock, buf, 512); // VULNERABILITY: Reads 512 bytes into a 64-byte buffer
printf("Processed connection data.\n");
}
int main(){
// Socket is basically the network channel
int sock = socket(AF_INET, SOCK_STREAM, 0), client; // 0 is for accepting the default transport protocol, which is TCP.
// htons just takes the port and converts it into a byte order that the network understands.
// sockaddress is the address for the network channel
struct sockaddr_in Sockaddress = {AF_INET, htons(31337), {0}}; // 0 is open, meaning listen on all my machine's IPs, not just one.
// Bind = attach the address to the network channel.
// sock is of course the socket network channel and we are binding it with the socket address.
// (struct sockaddr*) is a generic label bind accepts; the * makes it -> pointer to.
// &Sockaddress is the actual pointer to the sockaddr_in box we made earlier.
bind(sock, (struct sockaddr*)&Sockaddress, sizeof(Sockaddress));
listen(sock, 1); // Start accepting incoming connections on the socket & 1 = how many callers can wait on hold while you're busy with one.
while((client = accept(sock, 0, 0))){
// Remember we are the caller!
// client is whoever connects to your network channel. The first 0 is the caller's info or address, so we will not log that, and since the caller is us and we are on localhost there is no point saving that. The second 0 -> do not save the size of the IP address either.
// Let's call the handle_connection function.
handle_connection(client);
close(client);
}
}
Now recompile it -> gcc server.c -o server -fno-stack-protector -z execstack -no-pie -> then run python3 fuzz.py
server.c:13:5: warning: βreadβ writing 512 bytes into a region of size 64 overflows the destination [-Wstringop-overflow=]
13 | read(client_sock,buf, 512); // VULNERABILITY: Reads 512 bytes into a 64-byte buffer
| ^~~~~~~~~~~~~~~~~~~~~~~~~~
server.c:12:10: note: destination object βbufβ of size 64
12 | char buf[64];
| ^~~
In file included from server.c:3:
/usr/include/unistd.h:371:16: note: in a call to function βreadβ declared with attribute βaccess (write_only, 2, 3)β
371 | extern ssize_t read (int __fd, void *__buf, size_t __nbytes) __wur
| ^~~~
Processed connection data.
zsh: segmentation fault ./server
We will learn more, but our buffer overflow worked!
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.