ML functions
Functions in the ml package
ml.link_analysis
ml.link_analysisml.link_analysis(input: Link | URL, mode="default") -> LinkAnalysisOutput
LinkAnalysis analyzes a link and classifies it as benign or suspicious. The service sends suspicious URLs to a headless browser which resolves the effective URL and collects a screenshot. The screenshot is sent to an object detection model to detect brand logos, buttons, and input forms. We chose Phishpedia, an Open Source object detection project as our baseline model architecture.
If any logos are detected, those logos are cropped from the original screenshot and compared to a set of protected brand logos commonly used in credential phishing attacks. Discovered brands are available to MQL, along with summary information about login input boxes or captchas in the screenshot.
mode is an optional argument that alters LinkAnalysis's analysis criteria (see note below). By changing mode from its default of "default" to "aggressive", LinkAnalysis performs extra processing on a link when determining whether to fully analyze the link. For example, LinkAnalysis with mode="aggressive" will fetch the destination link of known common click trackers via HEAD and apply normal analysis criteria to that destination link.
View rules that use this function
// detect links to credential phishing pages
any(body.links,
all([ml.link_analysis(.)],
.credphish.disposition == "phishing"
and .credphish.brand.confidence in ("medium", "high")
)
)
// detect any links to credential phishing pages
any(body.links,
any([ml.link_analysis(., mode="aggressive")],
.credphish.disposition == "phishing"
and .credphish.brand.confidence in ("medium", "high")
)
)
// detect free subdomain links with a login or captcha
any(body.links,
all([ml.link_analysis(.)], (
.credphish.contains_login
or .credphish.contains_captcha
)
and (
.effective_url.domain.root_domain in $free_subdomain_hosts
or .original_url.domain.root_domain in $free_subdomain_hosts
))
)
// analyze the final DOM of a link within the body
any(body.links,
strings.icontains(ml.link_analysis(.).final_dom.display_text, "Redirect Notice")
and strings.contains(ml.link_analysis(.).final_dom.display_text, ".zip")
)
Analysis criteriaIn order to prevent LinkAnalysis from "clicking" on every link, such as Unsubscribes and one-time password resets, LinkAnalysis uses a URL classification model to determine which links to actually send to the service for analysis.
You can check whether LinkAnalysis
submitted,retrieved, oranalyzedthe target page by inspecting the response in the MQL editor.If you observe LinkAnalysis analyzing links it shouldn't or not analyzing links it should, consider whether the links you observe it clicking on are somewhat unique to your environment or are more specific to your email environment.
- You can configure Link Analysis domain and root domain exceptions in your environment by adding the domains you wish to be excluded to special Link Analysis exclusions.
- If you observe click behavior that you believe all Sublime users would benefit from, please send us an email or post in the Slack Community.
ml.logo_detect
ml.logo_detectml.logo_detect(input: File) -> [LogoDetectOutput]
LogoDetect uses computer vision to detect common brand logos used in attachment-based credential phishing attacks, such as impersonations of PayPal, Adobe, Microsoft, Outlook, Office365, DocuSign, and more. This includes embedded images in the body of messages as CIDs.
Our object detection model identifies logos, which are then cropped into separate images. These images are passed through a Siamese Neural Network to generate a feature vector. We compare this vector to a database of known logos using a similarity calculation. If the score exceeds a predetermined threshold, we confirm it as the respective brand logo.
For text-based logos, we utilize OCR, a computer vision technique for extracting text from images. Combined with Siamese Networks, this approach ensures comprehensive logo detection.
View Rules that use this function
// detect SharePoint logos in attached images
any(attachments,
.file_type in ('png', 'jpeg', 'jpg', 'bmp')
and any(ml.logo_detect(.).brands, .name == "Microsoft SharePoint")
)
// detect DocuSign logos in attached images
any(attachments,
.file_type in ('png', 'jpeg', 'jpg', 'bmp')
and any(ml.logo_detect(.).brands, .name == "DocuSign")
)
// detect Norton logos in attached PDFs
any(attachments,
.file_type == "pdf"
and any(ml.logo_detect(.).brands, .name == "Norton")
)List of Supported Brands
Don't see the brand you're looking for? Want to be able to detect your own company's logo? Contact your account team or send an email to [email protected] with a few examples of the logo and we'll get back to you.
ABN
ADP
AOL
AT&T
Adobe
AliExpress
Amazon
American Express
Apple
Awardco
BB&T Corporation
BBVA
BT
Bank of America
Barclays
Belastingdienst
Benteler
BeyondTrust
Bol
Box
CFA
CNA
CVS
Caixabank
Capital One Bank
CalPoly
Captcha
Carta
Chase
ChicagoTitle
Citi
Cloudflare
Coinbase
CyberArk
DHL
DKB
DPD
Dayforce
Digid
Discord
Discover
Disney
DocuSign
Dropbox
Ebay
Experian
Facebook
FakeAttachment
FanDuel
FedEx
FidelityTitle
FirstAm
FuboTV
GLS
GM
GeekSquad
Gemini Trust
Generic Captcha
Generic Webmail
Github
Gmail
GoDaddy
Google
GoogleDrive
Gusto
HSBC Bank
Heroku
Home Depot
HubSpot
Hulu
Huntress
ING
IRS
Instagram
Invite Company
JFrog
KPN
Kehe
Key Bank
LawyersTitle
Ledger
LinkedIn
Lloyds
M & T Bank
MadisonTitle
MailChimp
Mailgun
Mastercard
McAfee
Meta
MetaMask
Microsoft
Microsoft Office365
Microsoft OneDrive
Microsoft Outlook
Microsoft SharePoint
Microsoft Teams
NATO
NHS
NatWest
Navan
Navy Federal Credit Union
Netflix
Norton
OVO
Okta
OldRepublicTitle
OpenAI
PNC
Palo Alto Networks
Pandora
PayPal
PostNL
Postbank
Proton
Pulley
QuicklySign
Quickbooks
RBS
Rabobank
Rakuten
Robert Half
RoyalMail
SBB
SSA
Santander
Schwab
SendGrid
Shein
Signal
Silicon Valley Bank
Slack
Snowflake
Sparkasse
Spotify
Square
StewartTitle
Stratus
Stripe
SunTrust Bank
Swiss Post
Swisscom
TD Bank
Target
Targobank
Threads
TicorTitle
Tidal
TikTok
Trezor
TrustWallet
Tyrell
U.S. Bank
UCSB
UPS
USPS
Vanguard
Venmo
Visa
Vodafone
Volksbank
WeTransfer
Wells Fargo
Wex
WhatsApp
Wise
Workday
WoS
X
Yahoo
Zebra
Zelle
Zendesk
Ziggo
Zoom
Zscalerml.macro_classifier
ml.macro_classifierml.macro_classifier(input: File) -> MLMacrosOutput
The Sublime Macro Classifier introduces machine learning in MQL to detect malicious VBA macro attachments. Combining ML and MQL allows users to combine the model output with custom detection logic to surface what matters most while reducing the noise commonly associated with black-box ML approaches.
The classifier uses XGBoost to analyze VBA keywords, file metadata, and Oletools output to predict whether an attachment is likely to cause harm.
Use ml.macro_classifier to detect suspicious VBA macro attachments.
View rules that use this function
// detect malicious VBA macros in Office documents, high confidence
any(attachments, .file_extension in~ ("doc", "docm", "docx", "dot", "dotm", "pptm", "ppsm", "xlm", "xls", "xlsb", "xlsm", "xlt", "xltm", "zip")
and ml.macro_classifier(.).malicious
and ml.macro_classifier(.).confidence in ("high")
)
// detect malicious VBA macros in Office documents, low or medium confidence
any(attachments, .file_extension in~ ("doc", "docm", "docx", "dot", "dotm", "pptm", "ppsm", "xlm", "xls", "xlsb", "xlsm", "xlt", "xltm", "zip")
and ml.macro_classifier(.).malicious
and ml.macro_classifier(.).confidence in ("low", "medium")
)ml.nlu_classifier
ml.nlu_classifierml.nlu_classifier(input: string) -> NluResult
Natural Language Understanding, or NLU, provides users with a machine learning service to analyze text-based content. The service has three primary capabilities:
- Email Classification
- Named Entity Recognition
- Topic Recognition
Email Classification
The Email Classification component takes a body of text as input and provides Intents and/or Tags.
Intent
Intents are top-level categories describing common language attackers use to carry out phishing attacks.
| Name | Description |
|---|---|
bec | Emails containing urgent language about quick tasks from C-suite, HR, and Accounting Depts. |
callback_scam | Emails containing language about renewing/purchasing services such as tech support, antivirus, or cryptocurrency. |
cred_theft | Emails contain language urging users to visit a link leading to a realistic-looking portal that requires their credentials to log in. |
extortion | Emails meant to intimidate victims with threats of blackmail. |
steal_pii | Emails requesting updates to billing information, personal identification, and tax returns. |
job_scam | Deceptive emails disguised as employment offers to dupe students into divulging sensitive data or becoming unwitting accomplices in criminal or fraudulent schemes. |
Tags
Tags are subcategories that provide additional context for financial-themed phishing attacks. The service returns the following values:
| Name | Description |
|---|---|
invoice | These emails contain language about viewing invoices via links or attachments. |
payment | These emails contain language about ACH, EFT, or Wire payments. |
purchase_order | These emails contain language about Purchase Orders, Requests for Quotation. |
Example Usage
type.inbound
and any([body.plain.raw, body.html.inner_text],
any(ml.nlu_classifier(.).intents,
.name == "bec" and .confidence == "high"
)
)
// first-time sender
and (
(
sender.email.domain.root_domain in $free_email_providers
and sender.email.email not in $sender_emails
)
or (
sender.email.domain.root_domain not in $free_email_providers
and sender.email.domain.domain not in $sender_domains
)
)Entity Recognition
Named Entity Recognition (NER) identifies, tags, and extracts important keywords within a body of text. Users can leverage this output to determine if an email contains language commonly associated with urgency, requests, or financial matters. The available entities are listed below:
| Name | Description | Examples |
|---|---|---|
greeting | Token(s) that aid in the identification of the recipient | hello, dear |
financial | Token(s) containing financial details such as payments, bank accounts, or real estate transactions | wire, bank details, ACH payment |
org | Token(s) containing an organization name | Google, Microsoft |
recipient | Token(s) representing the recipient of the email. Either a name or a generic designator. | Jane Doe, all |
request | Token(s) asking the recipient to act on behalf of the sender | "I need you to", "please open" |
salutation | Token(s) signifying the end of the correspondence, aids in the identification of the sender | thanks, regards |
sender | Token(s) representing the sender of an email. Either a name or a generic designator. | Ms. Tyrell, IT Department |
urgency | Token(s) containing language meant to urge recipient to act immediately | ASAP, immediately |
Example Usage
type.inbound
and sender.display_name in~ $org_display_names
and any(ml.nlu_classifier(body.current_thread.text).entities, .name == "urgency")
and any(ml.nlu_classifier(body.current_thread.text).entities, .name == "request")
Topic Recognition
The Topic Classification component takes a body of text as input and provides topic classification for email content analysis. It analyzes message content to identify the primary topics and themes present in the email, helping to categorize and understand message intent.
ml.nlu_classifier(string, display_name=sender.display_name, subject=subject.subject) -> TopicResponse
// First parameter is required text to analyze
// Optional named parameters provide additional contextParameters
- Required text input (string) - The text content to analyze
display_name(optional) - Sender display name. Defaults to sender.display_name if not providedsubject(optional) - Subject text. Defaults to subject.subject if not provided
When analyzing message body content, use body.current_thread.text to ensure valid input. The function can also analyze OCR text or other content sources.
Supported Topics
The function can identify the following topics:
Business & Professional
"Financial Communications"- Banking, investments, bills, invoices, financial services"Legal and Compliance"- Legal matters, terms of service, privacy policies, compliance"Customer Service and Support"- Support tickets, inquiries, feedback requests"Professional and Career Development"- Job opportunities, training, industry insights"E-Signature"- Electronic document signing requests and updates"B2B Cold Outreach"- Business-to-business prospecting and outreach messages"Benefit Enrollment"- Employee benefits, insurance enrollment, HR communications
Technology & Security
"Security and Authentication"- Account security, password resets, 2FA, login alerts"Software and App Updates"- Software changes, new features, bug fixes"File Sharing and Cloud Services"- Shared files, storage notifications, collaboration"Secure Message"- Encrypted messaging and confidential communications
Communications & Notifications
"Newsletters and Digests"- Regular content compilations and updates"Reminders and Notifications"- Event/task reminders, calendar notifications"Out of Office and Automatic Replies"- Absence notifications, auto-responses"Bounce Back and Delivery Failure Notifications"- Failed email delivery notices"Voicemail Call and Missed Call Notifications"- Alerts for voicemails, calls, and missed call notifications
Marketing & Promotions
"Advertising and Promotions"- Marketing emails, sales, product launches"Events and Webinars"- Event invitations, RSVPs, online/offline gatherings"Travel and Transportation"- Trip planning, bookings, travel updates"Contact List Solicitation"- Requests to join mailing lists or contact databases
Commerce & Transactions
"Order Confirmations"- Purchase confirmations, order status updates"Shipping and Package"- Delivery tracking, shipping notifications, package updates
Public & Community
"Government Services"- Official government communications"Emergency Alerts"- Urgent notifications, weather, public safety"News and Current Events"- News updates and current affairs"Political Mail"- Campaign messages, political updates"Charity and Non-Profit"- Fundraising, volunteer opportunities"Environmental and Sustainability"- Updates on environmental initiatives and sustainability efforts"Acts of Violence"- Reports or alerts about violent incidents or threats
Health & Education
"Health and Wellness"- Medical appointments, health insurance, wellness"Educational and Research"- Learning materials, academic announcements
Entertainment & Social
"Entertainment and Sports"- Movies, music, games, sports updates"Social Media and Networking"- Social network notifications, connections"Romance"- Dating-related messages, relationship communications
Sensitive Content
"Sexually Explicit Messages"- Adult content and explicit material
Common Use Cases
Basic Usage
type.inbound
and any(ml.nlu_classifier(body.current_thread.text, display_name=sender.display_name, subject=subject.subject).topics,
.name == "Voicemail Call and Missed Call Notifications"
and .confidence == "high"
)Analyze Attachments with OCR
Combine with OCR functionality to analyze text in attachments:
type.inbound
and any(attachments,
.file_type in ('pdf', 'png', 'jpeg', 'jpg') and
any(file.explode(.),
any(ml.nlu_classifier(.scan.ocr.raw).topics,
.name == "Financial Communications"
and .confidence == "high"
)
)
)Negate Topics
Identify out of office auto replies, and email bounce back notification message types:
type.inbound
// your detection logic that you would like to exclude OOO replies from
and not any(ml.nlu_classifier(body.current_thread.text).topics,
.name in (
"Out of Office and Automatic Replies",
"Bounce Back and Delivery Failure Notifications"
)
and .confidence == "high"
)Multi-Topic Analysis
Check for multiple topics to further scope a query:
type.inbound
and any(ml.nlu_classifier(body.current_thread.text).topics,
.name in ("Political Mail")
and .confidence == "high"
)
and not any(ml.nlu_classifier(body.current_thread.text).topics,
.name in (
"Charity and Non-Profit",
"News and Current Events",
"Government Services"
)
and (.confidence == "high" or .confidence == "medium")
)Best Practices
- Confidence Levels
- Prefer
confidence == "high"when topics are critical to a rule or hunt - Consider medium confidence for supplementary signals
- Avoid using low confidence results in isolation
- Prefer
- Topic Combinations
- Consider both presence and absence of topics
- Use with other detection methods for scoping
- Performance
- Avoid unnecessary topic analysis on filtered messages
- Consider using other methods for simple text matching
- Input Selection
- Provide specific content for targeted analysis
- Consider context when analyzing extracted text
Considerations
It is important to remember that the NLU engine only looks at text. Because of this, it needs additional context to be an adequate detector. For example, attackers may craft an email that looks the same as a password reset for your favorite social network. The NLU engine would classify the text as cred_theft, but it would also do the same for a legitimate password reset email. But pairing it with a First-Time/Unsolicited Sender or LinkAnalysis provides the necessary context to make an effective detector.
Updated about 13 hours ago