Breaking Down Appsec Part 4: Following the Trail (Taint analysis)
Part of appsec and hacking is about understanding how to connect the dots
Continuing in our series in breaking down application security so that anyone can do it without hiring expensive teams of employees, we talk about how to understand what makes an app “hackable” through something called “Taint Analysis”.

Hack The Planet
So i’m not going to go into the entire history of the word hacking, but they can be read here, here, and here.
Basically, hackingmeant finding clever solutions to solving problems or creating stuff, but some maintain that the word should actually be “Cracker” for people who break stuff. Mainstream media took over with some of the movies in the 1980s (Wargames, Hackers, etc.) and the usage of the word was solidified. Most technologists have heard of “hackathons” which is really the usage of the original meaning (hack) + marathon. For the purposes of our blog, we’ll continue to use the mainstream definition of hacking, which is knowingly accessing a "protected computer" without authorization or exceeding authorized access to commit illegal acts (as stated by the CFAA).
BUT, what does hacking actually mean from the perspective of Application Security? We blend and borrow the historical and the modern definitions of “hacking”:
Finding clever solutions to knowingly achieve goals with malicious intent.
The definition is a bit abstract because goals themselves vary and ways to achieve these abstract goals also vary. I will provide an example to help elucidate this.
A note before we move on: with applications being deployed so that anyone in the world can access your application/site, you can be hacked from anywhere in the world as long as there is a network connection available.
Methodology
Quickly, before I provide examples, you have to know how these people are hacking your application.
They do this by application context. They may or may not already have your source code, fine. If they don’t, they start by exploring your application, carefully observing how workflows work by using your application as a regular user.
They also proxytraffic to observe all the detailed network requests that your browser is making through a software proxy (usually burpsuiteor something similar). Basically some software sits between a user’s browser and your application’s server and from that software, the user is able to change anything about the web request to your server and capture whatever is sent back to observe how your application’s server responds. To simplify it looks like below:

If none or only some of the above makes sense to you, I can break this down for you a bit further in another post delving into the networking fundamentals! Just let me know!
One Oz, Multiple Brick Roads (example)
I talked about examples of clever solutions with achieving malicious goals. We have to come up with a nice scenario for this to make this sensible.
Let’s assume we have an application (web/mobile/etc) for our restaurant called HackDonalds where we serve some fine foods like Burgers, Fries, and milkshakes. A hacker finds about this app/site and thinks “Wow, this sounds delicious! I want to explore what they have on their menu”. They do their exploration, but because they’re a hacker, they’re observing the ordering workflows, account creation, and proxying the web traffic, like we discussed above. The application seems quite complex!
Great the hacker now has some understanding of the application!
Now the hacker realizes, that he/she doesn’t really want to pay for the food, but would still like to get food. Okay, so then how does he/she do that?
Well, in order to get free food, the question is who pays for free food? There are 3 parties here who can pay for the food. The hacker, other users, or Hackdonalds.
We’ve broken up the problem of free food, to TWO smaller substeps that converge to the high level goal, meaning there isn’t just one “yellow brick road” to Oz, there’s actually 2! If one of them pay, then you don’t have to!
So now that the hacker knows there are 2 ways, the hacker can focus on his/her efforts to solving any one of these problems through different means.
There might be a POST request where they may be able to submit orders as a different person (authorization issue), or an open GET route to order food items as a staff/admin (authentication issue), or there could be remote code execution where the hacker can just force the application server to do whatever he/she wants. Each of these issues could have a step before them as well, so we define the next step to achieve these sub-sub-goals. Each step backwards grows the different ways to achieve our original goal, with each issue growing into other ways to achieve that sub-goal itself.
All of a sudden there are SO many options for the hacker to exploit flaws within how the application is written and setup. It went from 1 to 2 sub-goals to potentially 10s if not 100s of ways to achieve those 2 sub-goals. In security language, we can create an Attack Treeof different ways to achieve our original goal. We’ll talk more about Attack Trees and Threat modeling in a later post.
Connect The Dots
In the above section, we discussed how with a goal, we can break it down then work backwards and find that there could be multiple ways to achieve your high level goal (BONUS LIFE TIP: this same workflow applies to most things in life so take that as you will). We can take a similar perspective to understanding your application and why application security vulnerabilities exist in software, in general.
All of taint analysis and security vulnerabilities can be attributed to connecting the dots.

What does connecting dots have to do with anything? Well, if you actually complete the connect the dots puzzle above, you’ll see a dinosaur of some sorts. But how does this come to be?
You can see that by tracing in order from the number 1 to 44, it completes a drawing that you can then observe that it achieves an objective/goal, of drawing a figure.
Similarly, application security has something called “taint analysis” where we connect dots to see whether a vulnerability exists or not. Like connect the dots, we have clear start and end dots.
The start dots are what we call “sources”.
The end dots are what we call “sinks”.
Sources are anywhere that the application uses potential user input from a web request. These can be things like any part of the request such as request body (the JSON blobs that come in to your application), request parameters (the part after the URI e.g. https://google.com/search?q=blahblah&…..), cookies, headers, and anything else you might use from the web request to progress your business logic in your application. These are the entry points to your application that hackers will manipulate to see what your application will respond with. If a source is able to be manipulated from the web request, we call this untrusted source(s).
Sinks are where a consequence might happen. I say consequence because there are tons of things that can happen dependent on an application, but here are a few examples: database calls (SQL injection), authorization issues(Broken Access Control), server commands (remote command execution), running pieces of code (remote code execution), displaying an HTML view back to your users (cross-site scripting), and so much more. Things that seem harmless like database calls or reading a file from your server could be extremely harmful.
So, now that we have sources in your code, all we have to do is connect the dots to sinks! We track the usage of the sources in your code, see how the program evolves while using the source(s) and see if it reaches sinks uninhibited. I say uninhibited because those sources may be manipulated to be safe or not. If it’s safe, we don’t have to track it further than the code location of the program where it does become safe. If it’s not “safe”, then we must track how that piece of data evolves in the program until it reaches one of our sinks. Once we’ve connected the dots, we can be somewhat confident that there’s a sink with a potential consequence there.
Cracking The Code (example)
Here’s a code example that we can follow:

So ignore the sections where it says to ignore in black.
Focus on the blue section.
We have a working Python FastAPI application where we have:
a POST request that receives a web request with a body shaped in a certain way (the UserLookUp request schema is right above the POST request). The payload comes directly from the POST route (step 1 in green [THIS IS THE UNTRUSTED SOURCE], circled in red)
The payload is parsed for the username field in a string (step 2 in green, circled in red).
The string is assigned to a value called “query_string” (step 3 in green, circled in yellow). We now consider this “query_string” value that contains the untrusted source as untrusted. It’s (“query_string”) been tainted.
The tainted “query_string”, containing the untrusted source is then placed inside a text function which converts the query_string into a text string (step 4, circled in yellow).
The tainted “query_string” text is then placed into a db.execute function call (step 5, underlined in black), which if you’re not aware, is a very common function to execute a query in a database.
We’ve connected the POST request body (source) all the way to the db.execute (SQL sink) to understand that we have a SQL injection vulnerability. It’s called SQL injection because we can end the query that the code was written for, and inject a SQL query of our own to achieve everything from getting everything from the database, changing the database, deleting the database, or inserting whatever we want.
Now that you’ve understood how to do this, go out and find some of your own “hackable” vulnerabilities! We’ll talk about how to test these code vulnerabilities in future posts!
Pigeon is a NYC Cybersecurity Services company, specializing in Application Security. If you or anyone you know needs application security services, please reach out to me at david@pigeonlabs.ai. We’re willing to work with you near and far!
Subscribe to Breadcrumbs
New field notes on appsec and AI agent security. Free — unsubscribe anytime.
