Securing Your Law Firm logoSecuring Your Law Firm
Back to resources
Blog post2026-08-19By Securing Your Law Firm

The Bots Scanning Your Law Firm's Website Aren't Just Bots Anymore

AI agents now scan law firm websites and DNS records for weaknesses at machine speed. See what's actually out there right now, and what to check first.

Blog post - By Securing Your Law Firm

Get your free Proprietary Exposure Review | Browse services

Every law firm website gets scanned constantly. That part is not new – automated traffic has probed for outdated plugins, exposed logins, and forgotten subdomains for years. What changed over the past twelve months is what is doing the scanning, and how much of the decision-making now happens without a human at the keyboard.

In this article:

  • Why an AI system now ranks as a top bug bounty researcher, and what that means for public websites
  • How criminal groups weaponized an open-source AI attack framework within hours of its release
  • What two nation-state-linked incidents reveal about how automated the exploit pipeline has become
  • Why the "ambient" scanning traffic hitting every site online has shifted in the attacker's favor
  • Why this all starts at DNS and certificate records, not the website itself
  • A short checklist to see what an automated agent would find first

The scanners got a decision-making layer

An autonomous system called XBOW spent 2025 climbing the HackerOne leaderboard, eventually becoming the first AI to rank as the top bug bounty researcher in the United States – ahead of thousands of human testers. Its architecture runs a coordinator agent that maps a target and decides what deserves attention, then hands the work to specialized attack agents that probe different vulnerability classes in parallel. By late 2025, XBOW had turned that capability into a self-serve product priced from around $6,000, with results delivered in about five business days. The point is not that this tool is dangerous on its own. It is that pointing an autonomous pentester at any public website is now something anyone can buy off a menu.

The version of this without a purchase agreement is HexStrike-AI, an open-source framework built for red teams that orchestrates more than 150 offensive security tools through an AI reasoning layer. According to Check Point Research, criminal groups began discussing how to weaponize it within hours of its public release, and were later observed using it to exploit newly disclosed Citrix NetScaler vulnerabilities in under ten minutes – a process that previously took skilled attackers days. Forked versions with the human-in-the-loop safeguards stripped out have since circulated on dark web forums.

When the target is bigger than a bug bounty

In early November 2025, Google's Threat Intelligence Group reported the first documented cases of malware that calls an AI model mid-execution rather than relying on hard-coded instructions: one strain rewrites its own obfuscation code on the fly by querying Gemini's API, another uses an LLM to generate the exact commands it needs to steal documents and system information from an infected machine.

Days later, Anthropic disclosed that a Chinese state-sponsored group had manipulated its Claude Code tool into running an espionage campaign against roughly thirty organizations, with the AI handling an estimated 80 to 90 percent of the operational work – reconnaissance, exploit development, credential harvesting, and lateral movement – with human operators stepping in at only four to six decision points per target. Neither case targeted a small law firm specifically. Both show that the reconnaissance-to-exploit pipeline these groups now run is increasingly automated, which means the cost of aiming it at any given target, including yours, keeps dropping.

The ambient noise changed too

Even outside these headline cases, the baseline traffic hitting every website online has shifted. Imperva's 2026 Bad Bot Report found automated traffic accounted for more than 53 percent of all web traffic in 2025, up from 51 percent the year before – meaning human visitors are now the minority on the open internet. GreyNoise's 2026 State of the Edge Report, based on nearly three billion malicious sessions observed over a 162-day period, found that 52 percent of remote code execution attempts came from IP addresses with no prior history in its dataset at all, and tracked one credential-spraying botnet that grew from 2,000 to 300,000 participating IP addresses in just 72 days. In practice, that means the infrastructure scanning a firm's website or mail server today may not exist tomorrow, rotating faster than reputation-based blocklists can keep up.

Why this starts at DNS, not the website

Before any of these systems – commercial, criminal, or freelance – ever send a request your website would log, they read what is already public. Certificate transparency logs (the same public record every HTTPS certificate gets logged to) reveal every subdomain a firm has ever issued a certificate for, including staging sites, old client portals, and vendor integrations everyone forgot about. DNS records show whether SPF, DKIM, and DMARC are configured to actually block spoofed email or just to log it quietly. WHOIS and registrar data reveal whether a near-identical domain is sitting unregistered, waiting for someone else to claim it. None of this requires touching the firm's systems. It is the same public exposure surface our free Proprietary Exposure Review checks, and it is exactly where an automated agent starts building a profile of a target before deciding whether continuing is worth the effort.

From there, the website itself gets fingerprinted: CMS and plugin versions, exposed admin logins, TLS configuration, and increasingly, any AI or automation tooling the firm has bolted on. GreyNoise recorded more than 91,000 scanning sessions targeting exposed Ollama AI inference servers between October 2025 and January 2026 – including a single eleven-day campaign that accounted for over 80,000 of those sessions – which shows that new categories of internet-facing tooling get added to the scanning rotation almost as fast as firms adopt them.

Why size does not protect you

An autonomous agent reading DNS records does not know, or care, whether it is looking at a five-attorney firm or a five-hundred-attorney one. The signals look the same either way, and the scan takes the same few seconds. Law.com reporting found that as large firms invest heavily in cybersecurity, bad actors targeting the legal industry have increasingly turned to small and mid-size firms instead, where lower per-target payoffs are offset by a higher volume of successful attacks. In Connecticut alone, breach notifications obtained through a public records request show law firm breaches rose from about three in 2020 to ten in 2025, concentrated mostly at smaller and midsize firms.

Andrew Chase, a cybersecurity litigation partner at Constangy Brooks Smith & Prophete, told Law.com that social engineering targeting smaller firms' operating and trust accounts is now "one of the most prevalent, prolific types of incidents," and several experts in the same piece said a growing share of these scams are AI-assisted, with attackers using AI to impersonate internal email systems and firm personnel. Per the ABA's own 2023 Legal Technology Survey, 29 percent of law firms have experienced a security breach, only 34 percent have an incident response plan in place, and 19 percent do not even know whether they have been breached. Business email compromise alone accounted for $2.9 billion in reported losses to the FBI in 2023. None of those numbers assume a nation-state adversary – they assume the ordinary, automated background scanning that now touches every firm's public footprint, whether anyone at the firm is watching for it or not.

What to check before an agent finds it first

None of this requires a security team or a big budget – just five checks most firms have never run:

  • Pull your own certificate transparency history and confirm every subdomain it returns is one you recognize and still intend to run.
  • Check whether SPF, DKIM, and DMARC are set to actually reject or quarantine spoofed mail, not just log it.
  • Search for lookalike domains close to your own, and register or monitor the obvious variants.
  • Confirm your website's CMS, plugins, and TLS configuration are current, not just "working."
  • Get a baseline of what is actually visible from the outside before assuming no one is looking.

The bottom line

The scanning traffic hitting your firm's DNS and website today is not hypothetical, and it is no longer purely mechanical. Some share of it is already deciding, on its own, whether what it found is worth pursuing further. The firms in the best position are not the ones hoping to stay too small to notice. They are the ones who already know what that traffic can see – and our free Proprietary Exposure Review is the fastest way to find out before an autonomous agent does.