File functions
Functions in the file package
Files can be delivered via email in a variety of ways, including directly as an attachment or auto-downloaded via links.
file.explode
file.explodefile.explode(input: File | HTML) -> [FileExplodeOutput]
FileExplode uses Strelka, a file extraction and metadata collection system developed by Target.
Strelka uses a variety of scanners to parse files of a specific flavor and performs data collection and/or file extraction on them. Strelka can recursively extract nested files (like a Word doc within a Zip file), identify malicious scripts, suspicious executables and text, run analysis like OCR and Macro detection, and more. For more information on how Strelka works, see the official Strelka documentation.
For a list of all available scanners, see the Github repo or the official Strelka docs.
View detection rules that use this function
// detect HTML smuggling techniques
any(attachments, .file_extension in~ ('html', 'htm') and
any(file.explode(.), "unescape" in .scan.javascript.identifiers)
)
// detect encrypted zip files
any(attachments,
any(file.explode(.),
'encrypted_zip' in .flavors.yara
)
)
// detect attachments soliciting the user to enable macros using OCR
any(attachments,
any(file.explode(.),
strings.icontains(.scan.ocr.raw, "enable macros")
)
)
// detect macros with auto-open
any(attachments,
any(file.explode(.),
any(.scan.vba.auto_exec, . == "AutoOpen")
)
)
// detect macros calling an exe
any(attachments,
any(file.explode(.),
any(.scan.vba.hex, strings.ilike(., "*exe*"))
)
)
file.expand_archives
file.expand_archivesfile.expand_archives(input: File | HTML, max_depth=5) -> ExpandArchivesResult
The file.expand_archives function takes in a file and extracts nested files recursively up to a specified max_depth (input is not included in depth). file.expand_archives supports many common formats (e.g. zip, tar, rar, gz, bzip2) and will pass through non-archives. This is similar to the function file.explode, but instead of running scanners you can chain other MQL functions to the extracted files' content. The example expands all attachments and checks whether any of them are images that have a logo.
any(attachments,
any(file.expand_archives(.).files,
.file_type in $file_types_images
and any(ml.logo_detect(.).brands, .name == "DocuSign")
)
)file.html_screenshot
file.html_screenshotfile.html_screenshot(input: File) -> File
The file.html_screenshot function takes a screenshot of HTML files so that you can query the image. This allows you to run logo detect on HTML attachments — ml.logo_detect(file.html_screenshot(.)) — or send the result to file.explode, empowering you to run OCR and QR analysis.
any(attachments,
(
.file_extension in~ ("html", "htm", "shtml", "dhtml")
or .file_type == "html"
or .content_type == "text/html"
)
and any(ml.logo_detect(file.html_screenshot(.)).brands,
.name != null and .confidence in ("medium", "high")
)
)file.message_screenshot
file.message_screenshotfile.message_screenshot() -> File
The file.message_screenshot function takes a screenshot of the message using the message body's HTML section. This screenshot is the same as the one that shows in the Message Preview pane when viewing a message, and is a representation of what the end-user would see. The resulting file can be passed into other File analysis functions, such as file.explode or ml.logo_detect:
// Check for an embedded Microsoft logo
any(ml.logo_detect(file.message_screenshot()).brands,
.name == "Microsoft" and .confidence in ("medium", "high")
)
// Run OCR on a screenshot of the message
any(file.explode(file.message_screenshot()),
strings.ilike(.scan.ocr.raw, "*free cooler*")
)View detection rules that use this function
file.oletools
file.oletoolsfile.oletools(input: File) -> OleToolsOutput
Oletools, developed by Philippe Lagadec, analyzes Microsoft OLE2 files such as Microsoft Office documents for malware and other suspicious indicators.
Use file.oletools to analyze attachments for malware or suspicious indicators like VBA macros, remote OLE objects, encryption, and more.
View detection rules that use this function
// detect suspicious macros
any(attachments, file.oletools(.).indicators.vba_macros.exists)
any(attachments, file.oletools(.).indicators.vba_macros.risk == "high")
// detect potential attempts to exploit CVE-2021-40444 (https://msrc.microsoft.com/update-guide/vulnerability/CVE-2021-40444)
any(attachments, any(file.oletools(.).relationships, strings.ilike(.target, "*html:http*")))
// detect external OLE object relationships
any(attachments, file.oletools(.).indicators.external_relationships.count > 0)
// detect encrypted Office documents
any(attachments, file.oletools(.).indicators.encryption.exists)
// detect macros that attempt to auto-execute when the document is opened
any(attachments, any(file.oletools(.).macros.keywords, .type == "autoexec"))
// detect suspicious macro source code
any(attachments, strings.ilike(file.oletools(.).macros.vba_code_all_modules, "*kernel32*", "*GetProcessId*"))file.parse_eml
file.parse_emlfile.parse_eml(input: Attachment | File) -> MessageDataModel
The file.parse_eml function takes in an EML attachment (file extension .eml or content type message/rfc822) or file and parses it into an MDM.
any(attachments,
(.file_extension == "eml" or .content_type == "message/rfc822")
and strings.icontains(file.parse_eml(.).subject.subject, "invoice")
)file.parse_html
file.parse_htmlfile.parse_html(input: File) -> HTML
The file.parse_html function parses an HTML file from an attachment, returning the full raw along with display_text and inner_text. This empowers detections such as running NLU on the display_text and completing a regex on the HTML without custom scanners or YARA signatures.
any(attachments,
(
.file_extension in~ ("html", "htm", "shtml", "dhtml")
or .file_type == "html"
)
and regex.icontains(file.parse_html(.).raw,
"fromCharCode",
"charCodeAt",
"charAt",
"parseInt"
)
)file.parse_text
file.parse_textfile.parse_text(input: File) -> ParseTextOutput
The file.parse_text function parses a file from an attachment, returning the decoded text.
any(attachments,
(
.file_extension in~ ("html", "htm", "shtml", "dhtml")
or .file_type == "html"
)
and strings.icontains(file.parse_text(.).text, "invoice")
)Updated about 13 hours ago