October 1, 2026
My First Time Doing GitHub & Google Dorking on a Real Target
How I approached reconnaissance, what I actually found.

By Taoqui
9 min read
When I first started learning cybersecurity, I thought reconnaissance was mostly about running tools.
You know the workflow:
Tool โ Scan โ Results โ Find something interestingTool โ Scan โ Results โ Find something interestingBut for this exercise, I wanted to try something much simpler.
Google. GitHub. And a lot of curiosity.
For this write-up, my target was:
youtube.comyoutube.comI wasn't trying to break into YouTube. I wasn't exploiting anything.
I wanted to see what I could learn from information that was already publicly available.
Once I understood the basic idea of Google Dorking, I started experimenting with search operators.
At first, all the operators looked complicated.
After using them a few times, I realized they're basically just different ways of telling Google:
"Search here."_ "Search for this." _"Don't show me that."
Here's how I broke them down as a beginner.
Level 1 โ The Beginner's Toolkit
These are the operators I learned first.
site:
This was probably the most useful operator for my initial reconnaissance.
site:youtube.comsite:youtube.comInstead of searching the entire internet, I'm asking Google to focus on results associated with youtube.com.
I could then narrow the search further:
site:youtube.com helpsite:youtube.com helpor:
site:youtube.com privacysite:youtube.com privacyThis helped me understand what parts of the public website were being indexed.
intitle:
This searches for a word or phrase in the page title.
For example:
intitle:"YouTube"intitle:"YouTube"This can be useful when you're trying to understand how a website's pages are categorized.
For authorized security labs, you might also experiment with:
intitle:"login"intitle:"login"The important thing is not to treat every result as a vulnerability.
A login page is simply a login page.
Finding it doesn't mean you've found a security flaw.
inurl:
This looks for a keyword inside a URL.
For example:
site:youtube.com inurl:helpsite:youtube.com inurl:helpThis can help identify sections of a website based on their URL structure.
Again, the goal during passive reconnaissance is understanding the public footprint rather than attempting to access restricted functionality.
filetype:
This was another operator I found useful.
For example:
site:youtube.com filetype:pdfsite:youtube.com filetype:pdfThis asks Google to return PDF documents associated with the domain.
You can use the same concept with other public document formats when appropriate.
The interesting part isn't necessarily the document itself.
It's what the document tells you about the organization's public information footprint.
Quotation marks " "
Quotation marks force Google to look for an exact phrase.
For example:
"Welcome to YouTube""Welcome to YouTube"Instead of searching separately for each word, Google tries to find the phrase as a whole.
This becomes particularly useful when searching for a specific piece of wording.
-
The minus operator excludes something from the results.
For example:
site:youtube.com -shortssite:youtube.com -shortsThis can be useful when a particular type of result is overwhelming your search.
It's basically telling Google:
"Show me this, but leave that out."
Level 2 โ Combining Operators
This is where Google Dorking started becoming more interesting to me.
The individual operators aren't particularly complicated.
The real power comes from combining them.
For example:
site:youtube.com filetype:pdfsite:youtube.com filetype:pdfis much more specific than:
youtube pdfyoutube pdfThe first query gives Google two conditions:
- Focus on
youtube.com - Return PDF documents
That's the basic idea behind chaining operators.
Searching specific public content
For example:
site:youtube.com "privacy"site:youtube.com "privacy"This lets me investigate pages associated with YouTube that contain a particular phrase.
Another example:
site:youtube.com intitle:"YouTube"site:youtube.com intitle:"YouTube"Now I'm narrowing the search based on both the domain and the page title.
The more specific the question I'm asking, the more useful the results tend to become.
Level 3 โ Going Beyond Basic Searching
This is where I learned an important lesson.
There are plenty of advanced Google queries online that search for things such as:
- exposed credentials
- private keys
- database dumps
- configuration files
- backup files
- administrative interfaces
- exposed source-control directories
Technically, these queries can demonstrate interesting security concepts.
But I don't think beginners should immediately point them at real organizations.
Instead, use a deliberately vulnerable lab or a system you own.
For example, rather than searching the real internet for exposed credentials, create a local practice environment and experiment there.
That gives you the same learning experience without turning reconnaissance into unauthorized access.
One Important Lesson: A Dork Result Isn't Automatically a Vulnerability
This was something I had to remind myself of constantly.
Suppose a search returns:
login pagelogin pageThat doesn't automatically mean:
VULNERABILITYVULNERABILITYSuppose it returns:
PDF documentPDF documentThat doesn't automatically mean:
SENSITIVE DATASENSITIVE DATAAnd suppose GitHub contains:
API endpointAPI endpointThat doesn't automatically mean:
EXPLOITABLE APIEXPLOITABLE APIA Google search result is simply evidence that something is publicly indexed.
The next step is understanding what that information actually means.
The Google Hacking Database
Eventually I came across the Google Hacking Database (GHDB).
This was interesting because instead of inventing search queries from scratch, GHDB provides a large collection of categorized search queries used in security research.
It's maintained by Exploit Database and contains categories covering things such as:
- Files containing usernames
- Sensitive directories
- Web server detection
- Vulnerable files
- Footholds
- Various forms of information exposure
It's a useful resource for learning how search operators can be combined.
But there's an important distinction:
A query being listed in a security database doesn't mean you should run it against an arbitrary organization.
For real targets, authorization matters.
For learning, a lab is the safest place to experiment.
What I Actually Took Away From This
When I first saw lists of Google Dorks, I thought cybersecurity was about memorizing hundreds of complicated queries.
It isn't.
At least, that's not how I think about it anymore.
The operators are just building blocks.
site:
intitle:
inurl:
filetype:
intext:
"exact phrase"
-site:
intitle:
inurl:
filetype:
intext:
"exact phrase"
-The real skill is knowing what question you're trying to answer.
For my YouTube reconnaissance exercise, I wasn't trying to find a magic query that would suddenly reveal some secret.
I was trying to understand the public footprint.
That's a much more useful mindset.
My Beginner Dorking Checklist
When I'm doing authorized passive reconnaissance now, I think about it in roughly this order:
1. Identify the domain
โ
2. Search normally
โ
3. Use site:
โ
4. Narrow with keywords
โ
5. Investigate public documents
โ
6. Check relevant public repositories
โ
7. Verify what I find
โ
8. Document everything1. Identify the domain
โ
2. Search normally
โ
3. Use site:
โ
4. Narrow with keywords
โ
5. Investigate public documents
โ
6. Check relevant public repositories
โ
7. Verify what I find
โ
8. Document everythingAnd most importantly:
PUBLIC โ AUTHORIZED TO EXPLOITPUBLIC โ AUTHORIZED TO EXPLOITThat's probably the biggest lesson I learned from this entire exercise.
Google can help you discover information.GitHub can help you understand public code.
But what you do with that information is where responsible security research begins.
Understanding GitHub Dorking
After experimenting with Google Dorking, I moved on to another place that contains an incredible amount of publicly available technical information:
GitHub.
At first, I thought GitHub Dorking simply meant searching GitHub for passwords.
It's actually much broader than that.
You can use GitHub's search functionality to investigate repositories, organizations, file names, programming languages, paths, and development history.
I started thinking of it as digital archaeology.
You're not necessarily looking for a vulnerability. You're trying to understand what has been publicly published and what that information tells you.
1. What is GitHub Dorking?
GitHub Dorking is the practice of using GitHub's search syntax and filters to locate specific information within publicly accessible repositories.
For example, instead of searching GitHub for:
youtubeyoutubeyou can make your search much more specific.
You might search for:
org:googleorg:googleor:
language:pythonlanguage:pythonor:
extension:jsonextension:jsonThe difference is huge.
You're no longer asking:
"What does GitHub know about this?"
You're asking a much more specific question.
2. The GitHub Search Operators I Learned
These were some of the most useful operators I came across.
org:
Limits the search to repositories belonging to a particular GitHub organization.
For example, when researching an organization you are authorized to study:
org:example-orgorg:example-orgThis is useful because large organizations can have hundreds or thousands of repositories.
Rather than searching all of GitHub, you're narrowing the scope.
repo:
This limits your search to a particular repository.
For example:
repo:owner/projectrepo:owner/projectThis becomes useful once you've already identified a repository you want to understand.
You can then investigate specific files, technologies, or terms within that project.
filename:
This searches for a particular filename.
For example:
filename:README.mdfilename:README.mdor:
filename:docker-compose.ymlfilename:docker-compose.ymlor:
filename:package.jsonfilename:package.jsonThese searches can be useful for understanding the structure and technologies used by public projects.
extension:
This lets you narrow results by file extension.
For example:
extension:jsonextension:jsonor:
extension:yamlextension:yamlor:
extension:pyextension:pyYou can combine this with other search terms.
For example:
org:example-org extension:pyorg:example-org extension:pyNow you're asking GitHub for Python files associated with that organization.
path:
The path: qualifier is useful when you're interested in a particular directory.
For example:
path:docspath:docsThis can help when exploring the documentation structure of a public repository.
language:
This is one of my favorite filters because it gives you a quick idea of the technology involved.
For example:
language:pythonlanguage:pythonor:
language:javascriptlanguage:javascriptor:
language:javalanguage:javaYou can combine it with a repository or organization:
org:example-org language:pythonorg:example-org language:pythonNow the search is much more focused.
3. Combining Operators
This is where GitHub search becomes much more interesting.
You can combine multiple qualifiers.
For example:
org:example-org language:python filename:requirements.txtorg:example-org language:python filename:requirements.txtNow we're effectively asking:
"Show me Python dependency files belonging to this organization."
Another example:
repo:owner/project extension:jsonrepo:owner/project extension:jsonThis asks GitHub to focus on JSON files inside a particular repository.
The basic idea is:
Target
+
File
+
Language
+
Path
+
KeywordTarget
+
File
+
Language
+
Path
+
KeywordThe more clearly you define the question, the less noise you get.
4. Searching Configuration Without Hunting Credentials
Configuration files are interesting during reconnaissance because they can tell you a lot about an application's architecture.
For example:
filename:package.jsonfilename:package.jsoncan tell you about JavaScript dependencies.
Similarly:
filename:requirements.txtfilename:requirements.txtcan reveal Python dependencies.
And:
filename:docker-compose.ymlfilename:docker-compose.ymlcan provide clues about how a development environment is structured.
This is the kind of information I was interested in during my reconnaissance.
I wasn't trying to obtain passwords.
I wanted to understand:
"What technologies are being used?"
5. Looking at Development History
GitHub becomes even more interesting when you stop looking only at the current files.
Repositories also have development history.
You can investigate things such as:
- Commit dates
- Contributors
- Changes to files
- Project activity
- Previous versions of documentation
This is useful because the current repository is only one snapshot of a project's history.
Something that existed months ago may no longer exist today.
That doesn't automatically make the old information sensitive or exploitable, but it can provide useful historical context.
6. Temporal Searches
GitHub provides date-related search qualifiers that can help narrow results by repository activity.
For example:
org:example-org pushed:>2025-01-01org:example-org pushed:>2025-01-01This asks GitHub to focus on repositories pushed after a particular date.
You can use this when trying to understand which public projects are currently active.
The important thing I learned here is that time matters.
A repository that hasn't changed in years tells a very different story from one that is actively being developed.
7. What About Secrets?
This is where things get serious.
You will find plenty of articles online containing searches for:
API keys
passwords
cloud credentials
private keys
database credentials
tokensAPI keys
passwords
cloud credentials
private keys
database credentials
tokensThose searches demonstrate a real security problem:
Developers sometimes accidentally publish secrets.
But I don't think the correct lesson for a beginner is:
"Find a key and see whether it works."
The correct lesson is:
"Understand how accidental secret exposure happens and learn how to prevent it."
If you're practicing, use your own repository or a deliberately vulnerable security lab.
If you discover a possible secret in an authorized assessment, don't use it. Redact it from your notes and follow the applicable disclosure process.
8. A Safer Way to Practice
Instead of searching the public internet for somebody else's credentials, I created the concept of a small test repository.
For example:
my-security-lab/
โ
โโโ README.md
โโโ src/
โโโ docs/
โโโ config/
โโโ package.json
โโโ docker-compose.ymlmy-security-lab/
โ
โโโ README.md
โโโ src/
โโโ docs/
โโโ config/
โโโ package.json
โโโ docker-compose.ymlThen I can experiment with GitHub search against information I intentionally published.
For example:
repo:myusername/my-security-lab filename:README.mdrepo:myusername/my-security-lab filename:README.mdor:
repo:myusername/my-security-lab extension:jsonrepo:myusername/my-security-lab extension:jsonThis gives me the same opportunity to learn the search syntax without going after someone else's private information.
9. What I Actually Looked For During My YouTube Research
When I switched from generic GitHub searching to researching the public ecosystem around YouTube and Google, I focused on things that could help me understand the technology rather than attempting to obtain secrets.
My notes looked roughly like this:
SearchWhat I was looking foryoutubeGeneral GitHub referencesorg:Projects belonging to a known organizationlanguage:Programming languagesfilename:Important project filesextension:Particular types of source/configuration filespath:Specific directoriesrepo:Information within a particular projectpushed:Recently active repositories
This was enough to start building a picture.
10. The Biggest Lesson
The biggest mistake I made initially was thinking:
"A good security researcher must know hundreds of dorks."
I don't think that's true.
The syntax is useful, but the question you're asking is more important.
For example:
Bad approach:
"Give me secrets""Give me secrets"Better approach:
"What public projects are associated with this organization?""What public projects are associated with this organization?"Then:
"What technologies do those projects use?""What technologies do those projects use?"Then:
"What documentation and configuration information is publicly available?""What documentation and configuration information is publicly available?"Then:
"Is the information actually associated with the organization, or is it a third-party project?""Is the information actually associated with the organization, or is it a third-party project?"That progression is what makes reconnaissance useful.
11. Google + GitHub Together
This is where the two techniques started making sense to me.
Google helped me understand the public web footprint.
GitHub helped me understand the public development footprint.
I could think about the process like this:
TARGET
โ
โโโโโโโโโโดโโโโโโโโโ
โ โ
GOOGLE GITHUB
โ โ
Web pages Repositories
Documents Source code
Public info Documentation
โ โ
โโโโโโโโโโฌโโโโโโโโโ
โ
Recon notes
โ
Verified factsTARGET
โ
โโโโโโโโโโดโโโโโโโโโ
โ โ
GOOGLE GITHUB
โ โ
Web pages Repositories
Documents Source code
Public info Documentation
โ โ
โโโโโโโโโโฌโโโโโโโโโ
โ
Recon notes
โ
Verified factsNeither source gives you the entire picture.
But together, they can provide useful context.
Final Takeaway
GitHub Dorking sounded intimidating when I first encountered it.
After actually using GitHub's search functionality, I realized that most of the complexity comes from combining simple filters.
The operators are just building blocks:
org:
repo:
filename:
extension:
path:
language:
pushed:
created:org:
repo:
filename:
extension:
path:
language:
pushed:
created:The real skill is knowing what you're trying to learn.
And there's another lesson I don't want to skip:
Just because GitHub makes something searchable doesn't mean you're authorized to use it. Public information should be treated responsibly. For me, the goal of this exercise wasn't to collect passwords or tokens. It was to learn how much technical context can be discovered simply by asking better questions. And that turned out to be much more valuable.