SANDS Lab collects and analyzes real cyber threat data, generating the feature information and threat labels AI needs to learn, and builds it into a verifiable dataset. Drawing on years of experience running public-sector cybersecurity AI dataset projects and our own threat intelligence and AI technology, we handle the entire process — from building the data to validating and using the model.
Security AI performance starts with the threats it learns from.
Cyberattacks are constantly changing, with new malware, attack infrastructure, and techniques emerging all the time. For security AI to detect real threats, it needs training data with recency and context, accurate labels, and verified feature information — not just a large volume of data.
The latest threats must be continuously reflected.
Past data alone makes it hard to sufficiently learn newly emerging attack campaigns and variant threats.
A single IoC can't explain an attack.
You need interconnected context beyond hashes, IPs, domains, and URLs — attack groups, techniques, behavior, vulnerabilities, and more.
Wrong labels teach the AI wrong lessons.
Security data needs accurate refinement and labeling — not just malicious/benign status, but threat type and feature information too.
Data must be validated by whether a model actually learns from it.
Building a lot of data isn't the end of the process — a model's quantitative performance and real-world applicability must be validated too.
4 years of experience building cybersecurity AI datasets.
From 2021 to 2024, SANDS Lab has run public-sector cybersecurity AI dataset projects for four consecutive years, building data collection, processing, labeling, and validation systems across eight projects covering a range of cyber threat domains.
Cybersecurity AI Dataset R&D
Began public-sector cybersecurity AI dataset projects, building the data collection and processing system.
Cybersecurity AI Dataset R&D
Expanded the data-building domains and advanced the labeling and quality-validation process.
AI dataset refresh & advancement
Reflected the latest threat data to refresh the dataset and advanced the build system.
Strengthening AI dataset build & use
Strengthened the AI model validation and usage system for the datasets built.
Four years of experience across eight security data domains.
Malware
Experience building datasets based on malware file and behavior analysis.
Incident response
Experience building threat data accumulated through incident response.
Application security
Experience building application security vulnerability and threat data.
Proactive security monitoring
Experience building threat detection data for proactive security monitoring environments.
Threat profiling
Experience building attack group and campaign profiling data.
Threat hunting
Experience building data secured through threat hunting.
Threat intelligence
Experience building data based on threat intelligence analysis.
Recent incidents
Experience building data based on recent incident response cases.
Not just a file — we turn attack context into data.
Beyond simply collecting files and IoCs, SANDS Lab compiles the range of security feature information AI needs to learn and explain a threat — static and dynamic analysis, attack groups, attack techniques, vulnerabilities, and network-related information.
Malware & files
Compiles static and dynamic analysis information by file type — PE/EXE, PDF, ELF, APK — along with file structure and script/object/API call information.
Network IoCs
Compiles IPs, domains, URLs, DNS, Whois, along with related binaries and C2/distribution infrastructure information.
Threat context
Connects attack groups, campaigns, MITRE ATT&CK TTPs, vulnerabilities (CVEs), attack history, and threat type.
Incidents & intelligence
Compiles incident information, threat reports, OSINT, threat intelligence, and attack campaign context together.
We build across the entire data lifecycle —from collection to use.
Collect the latest threat data
Continuously collects the latest threat information — files, IoCs, reports, and more — using our own threat intelligence and a range of security information channels.
Static/dynamic analysis & feature extraction
Analyzes the collected raw threat data to generate feature information AI can learn from — file structure, behavior, network-related information, and more.
Label threat type & context
Labels threat information — malware type, attack group, attack technique — according to standardized criteria.
Validate data quality
Validates the quality of the build process and the resulting data against criteria such as readiness, completeness, usefulness, fitness, and accuracy.
Store & manage in a standardized format
Stores the dataset using structured formats such as JSON and STIX, in a form that can integrate with other security systems.
Confirm the AI actually learns
Trains an AI model on the built dataset and validates its effectiveness through quantitative metrics such as precision, recall, and F1 score.
Demonstrate, share & use
Applies the validated dataset to real security models and target environments, building a structure for use through search, API, packaging, and more.
Raw threat data, structured for AI.
Raw threat information, turned into a structure AI can read.
Transforms raw threat data — files, IPs, domains, URLs, reports — through static, dynamic, and contextual analysis into metadata that feeds into feature extraction.
Labels threat type down to the attacker.
Labels by threat type, attack group, and attack technique (TTP), structuring the data so AI can distinguish and explain threats.
A standardized data structure that connects to security systems.
Uses JSON for internal processing and system integration, and the international standard STIX 2.1 structure for cyber threat intelligence interoperability.
What matters more than volume is data a model can actually learn from.
The quality of a security AI dataset can't be judged simply by the volume of data collected. SANDS Lab validates it in stages — the quality of the build process, the accuracy of the data itself, AI model training results, and real-world applicability.
Readiness
Checks the foundation for data quality management — policy, regulations, organization, procedures.
Completeness
Checks that the defined data structure and input scope are fully covered without gaps.
Usefulness
Checks whether the data meets the actual AI training purpose and the needs of the people who'll use it.
Fitness
Reviews the data's fitness as training data — diversity, reliability, sufficiency, factuality, and more.
Accuracy
Validates the accuracy of ground truth, label, and metadata values.
We validate the data with a real AI model.
To confirm a built dataset is actually effective for security AI, we measure model performance during training and testing, then feed the detection results back into improving data quality.
File-based malware detection
Analyzes maliciousness based on file characteristics such as PE/EXE, PDF, and ELF.
APT attack group identification
Analyzes attack group and campaign context based on the characteristics of malware and attack behavior.
Malicious domain detection
Detects malicious infrastructure such as phishing and C2 using domains and related context.
The data proves its value through validated performance.
Multi-year model performance validation with external institutions
We trained AI models on the data built through our 2021–2024 public-sector cybersecurity AI dataset projects, and carried out quantitative performance validation against external institutional standards.
Explaining detection results in natural language
We also examined a structure that connects the AI model's detection results and related feature information with generative AI (LLM) to explain and summarize the reasoning behind a threat judgment in natural language.
So the data we build can power real-world security technology.

Security AI training & advancement
Used to train and improve the performance of security AI models covering malware, attack groups, malicious domains, and more.

Security model performance validation
Uses the validated dataset to quantitatively evaluate an AI model's detection performance and effectiveness.

Threat intelligence analysis
Connects IoCs with attack groups, campaigns, and TTPs, enabling analysis of threat context far richer than a simple indicator.

Corporate & institutional security R&D
Can be used as foundational data for research into new AI security models, threat analysis technology, and automation.
We build the data, validate it with AI, and feed it back into security technology.
A cybersecurity specialist, not a data vendor
We're not a general data company — we're a security company that has analyzed real cyber threats.
Built on our own threat data
We can build data on the foundation of our own threat intelligence platform and analysis data, backed by large-scale proprietary threat profiling data.
AI model R&D capability
We have the R&D capability to directly train and validate AI models on the data we build.
End-to-end experience
We have experience running the entire process — collection, processing, labeling, validation, storage, and use.
Frequently asked questions about Cybersecurity AI Dataset.
How is this different from general AI training data?
Cybersecurity AI data isn't just text or images — the relationships and context between threats matter, including malware, IoCs, attack groups, attack techniques, and vulnerabilities. SANDS Lab structures feature information and labels that security AI can learn from, based on real threat analysis technology.
What cybersecurity domains have you built data for?
Through public-sector projects from 2021 to 2024, we've built datasets across 8 domains: malware, incident response, application security, proactive security monitoring, threat profiling, threat hunting, threat intelligence, and recent incidents.
How do you validate data quality?
We inspect data quality at the collection, processing, and labeling stages, going through error analysis and improvement procedures. We then validate the data's effectiveness by training a real AI model on the built dataset and checking performance metrics such as precision, recall, and F1 score.
Do you support standard formats like STIX?
In our public-sector projects, we structured threat data using structured formats such as JSON and STIX 2.1, and built a data structure designed with interoperability with external security systems in mind.
Can you build a new dataset for a specific purpose?
SANDS Lab has experience building the entire data pipeline — from raw threat data collection through analysis, metadata generation, labeling, quality validation, and AI model validation. The actual build scope is designed based on the research purpose, data rights, and the required threat domain.
Can I purchase the dataset described on this page directly?
This page is an R&D page introducing SANDS Lab's cybersecurity AI data building and validation capability, rather than a specific commercial dataset product. The scope of data provision, joint research, or build collaboration is discussed separately based on data rights and project terms.
Let’s design the data your security AI needs —together.
From cyber threat data collection, processing, and labeling to quality validation and AI model demonstration — put SANDS Lab's data and AI research capability to work.
How does this product behave in your environment?
Whether you're exploring, evaluating, or rolling out, you connect directly with a SANDS Lab solutions engineer. Clear every question before contract — that's the point.
Product evaluation & PoC
Real-data PoCs, technical deep-dive sessions, and custom integration scoping. Everything you'd need to validate technical fit before the paperwork starts.
Technical collaboration & licensing
If you want to use the product in an academic benchmark or co-authored paper, we support research licenses and the underlying datasets. Co-authorship is on the table.