Skip to content
CYBERSECURITY AI DATASET

SANDS Lab collects and analyzes real cyber threat data, generating the feature information and threat labels AI needs to learn, and builds it into a verifiable dataset. Drawing on years of experience running public-sector cybersecurity AI dataset projects and our own threat intelligence and AI technology, we handle the entire process — from building the data to validating and using the model.

4 YEARS — Consecutive public-sector cybersecurity dataset projects8 DATASET DOMAINS — Experience building cybersecurity dataEND-TO-END — From collection to quality validation and useAI-VALIDATED — Data effectiveness verified with AI models
WHY SECURITY DATA MATTERS

Security AI performance starts with the threats it learns from.

Cyberattacks are constantly changing, with new malware, attack infrastructure, and techniques emerging all the time. For security AI to detect real threats, it needs training data with recency and context, accurate labels, and verified feature information — not just a large volume of data.

01FRESHNESS

The latest threats must be continuously reflected.

Past data alone makes it hard to sufficiently learn newly emerging attack campaigns and variant threats.

02CONTEXT

A single IoC can't explain an attack.

You need interconnected context beyond hashes, IPs, domains, and URLs — attack groups, techniques, behavior, vulnerabilities, and more.

03LABEL QUALITY

Wrong labels teach the AI wrong lessons.

Security data needs accurate refinement and labeling — not just malicious/benign status, but threat type and feature information too.

04VALIDATION

Data must be validated by whether a model actually learns from it.

Building a lot of data isn't the end of the process — a model's quantitative performance and real-world applicability must be validated too.

PROVEN R&D EXPERIENCE

4 years of experience building cybersecurity AI datasets.

From 2021 to 2024, SANDS Lab has run public-sector cybersecurity AI dataset projects for four consecutive years, building data collection, processing, labeling, and validation systems across eight projects covering a range of cyber threat domains.

012021

Cybersecurity AI Dataset R&D

Began public-sector cybersecurity AI dataset projects, building the data collection and processing system.

022022

Cybersecurity AI Dataset R&D

Expanded the data-building domains and advanced the labeling and quality-validation process.

032023

AI dataset refresh & advancement

Reflected the latest threat data to refresh the dataset and advanced the build system.

042024

Strengthening AI dataset build & use

Strengthened the AI model validation and usage system for the datasets built.

8 DATASET DOMAINS

Four years of experience across eight security data domains.

01

Malware

Experience building datasets based on malware file and behavior analysis.

02

Incident response

Experience building threat data accumulated through incident response.

03

Application security

Experience building application security vulnerability and threat data.

04

Proactive security monitoring

Experience building threat detection data for proactive security monitoring environments.

05

Threat profiling

Experience building attack group and campaign profiling data.

06

Threat hunting

Experience building data secured through threat hunting.

07

Threat intelligence

Experience building data based on threat intelligence analysis.

08

Recent incidents

Experience building data based on recent incident response cases.

SECURITY DATA COVERAGE

Not just a file — we turn attack context into data.

Beyond simply collecting files and IoCs, SANDS Lab compiles the range of security feature information AI needs to learn and explain a threat — static and dynamic analysis, attack groups, attack techniques, vulnerabilities, and network-related information.

01 — MALWARE & FILE

Malware & files

Compiles static and dynamic analysis information by file type — PE/EXE, PDF, ELF, APK — along with file structure and script/object/API call information.

02 — NETWORK IoC

Network IoCs

Compiles IPs, domains, URLs, DNS, Whois, along with related binaries and C2/distribution infrastructure information.

03 — THREAT CONTEXT

Threat context

Connects attack groups, campaigns, MITRE ATT&CK TTPs, vulnerabilities (CVEs), attack history, and threat type.

04 — INCIDENT & INTELLIGENCE

Incidents & intelligence

Compiles incident information, threat reports, OSINT, threat intelligence, and attack campaign context together.

END-TO-END DATA PIPELINE

We build across the entire data lifecycle —from collection to use.

Structured Dataset Generation PipelineDiverseCybersecurityDatasetsMalwareThreat ProfilingIncidentsActive MonitoringThreatIntelligenceApplicationSecurityThreat HuntingLatest IncidentsVerified Datasetfor AI UseOptimized for AITrainingHigh-quality DataKept Up to DateField-testedDataset01CollectDiverse threatsources02Clean &NormalizeDeduplication03LabelingExpertclassification04ValidationDouble review05ContinuousUpdateReflects latestthreatsContinuous Generation of Fresh Threat DatasetsScheduled GenerationRegular, automated dataset generationLabel VerificationExpert-driven label verificationQuality ControlOngoing quality monitoring & improvement
0101 COLLECT

Collect the latest threat data

Continuously collects the latest threat information — files, IoCs, reports, and more — using our own threat intelligence and a range of security information channels.

0202 ANALYZE & PROCESS

Static/dynamic analysis & feature extraction

Analyzes the collected raw threat data to generate feature information AI can learn from — file structure, behavior, network-related information, and more.

0303 LABEL

Label threat type & context

Labels threat information — malware type, attack group, attack technique — according to standardized criteria.

0404 VALIDATE

Validate data quality

Validates the quality of the build process and the resulting data against criteria such as readiness, completeness, usefulness, fitness, and accuracy.

0505 STORE

Store & manage in a standardized format

Stores the dataset using structured formats such as JSON and STIX, in a form that can integrate with other security systems.

0606 MODEL VALIDATION

Confirm the AI actually learns

Trains an AI model on the built dataset and validates its effectiveness through quantitative metrics such as precision, recall, and F1 score.

0707 UTILIZE

Demonstrate, share & use

Applies the validated dataset to real security models and target environments, building a structure for use through search, API, packaging, and more.

AI-READY DATA

Raw threat data, structured for AI.

01

Raw threat information, turned into a structure AI can read.

Transforms raw threat data — files, IPs, domains, URLs, reports — through static, dynamic, and contextual analysis into metadata that feeds into feature extraction.

02

Labels threat type down to the attacker.

Labels by threat type, attack group, and attack technique (TTP), structuring the data so AI can distinguish and explain threats.

03

A standardized data structure that connects to security systems.

Uses JSON for internal processing and system integration, and the international standard STIX 2.1 structure for cyber threat intelligence interoperability.

QUALITY & VALIDATION

What matters more than volume is data a model can actually learn from.

The quality of a security AI dataset can't be judged simply by the volume of data collected. SANDS Lab validates it in stages — the quality of the build process, the accuracy of the data itself, AI model training results, and real-world applicability.

READINESS

Readiness

Checks the foundation for data quality management — policy, regulations, organization, procedures.

COMPLETENESS

Completeness

Checks that the defined data structure and input scope are fully covered without gaps.

USEFULNESS

Usefulness

Checks whether the data meets the actual AI training purpose and the needs of the people who'll use it.

FITNESS

Fitness

Reviews the data's fitness as training data — diversity, reliability, sufficiency, factuality, and more.

ACCURACY

Accuracy

Validates the accuracy of ground truth, label, and metadata values.

MODEL-BASED VALIDATION

We validate the data with a real AI model.

To confirm a built dataset is actually effective for security AI, we measure model performance during training and testing, then feed the detection results back into improving data quality.

FILE MALWARE DETECTION

File-based malware detection

Analyzes maliciousness based on file characteristics such as PE/EXE, PDF, and ELF.

APT ATTRIBUTION

APT attack group identification

Analyzes attack group and campaign context based on the characteristics of malware and attack behavior.

MALICIOUS DOMAIN DETECTION

Malicious domain detection

Detects malicious infrastructure such as phishing and C2 using domains and related context.

VALIDATION RECORD

The data proves its value through validated performance.

Multi-year model performance validation with external institutions

We trained AI models on the data built through our 2021–2024 public-sector cybersecurity AI dataset projects, and carried out quantitative performance validation against external institutional standards.

Explaining detection results in natural language

We also examined a structure that connects the AI model's detection results and related feature information with generative AI (LLM) to explain and summarize the reasoning behind a threat judgment in natural language.

DATA UTILIZATION

So the data we build can power real-world security technology.

Security AI training & advancement
AI MODEL TRAINING

Security AI training & advancement

Used to train and improve the performance of security AI models covering malware, attack groups, malicious domains, and more.

Security model performance validation
MODEL VALIDATION

Security model performance validation

Uses the validated dataset to quantitatively evaluate an AI model's detection performance and effectiveness.

Threat intelligence analysis
THREAT INTELLIGENCE

Threat intelligence analysis

Connects IoCs with attack groups, campaigns, and TTPs, enabling analysis of threat context far richer than a simple indicator.

Corporate & institutional security R&D
SECURITY R&D

Corporate & institutional security R&D

Can be used as foundational data for research into new AI security models, threat analysis technology, and automation.

AI R&D CAPABILITY

We build the data, validate it with AI, and feed it back into security technology.

01 — CYBERSECURITY DOMAIN

A cybersecurity specialist, not a data vendor

We're not a general data company — we're a security company that has analyzed real cyber threats.

02 — PROPRIETARY DATA FOUNDATION

Built on our own threat data

We can build data on the foundation of our own threat intelligence platform and analysis data, backed by large-scale proprietary threat profiling data.

03 — AI MODEL ENGINEERING

AI model R&D capability

We have the R&D capability to directly train and validate AI models on the data we build.

04 — END-TO-END EXPERIENCE

End-to-end experience

We have experience running the entire process — collection, processing, labeling, validation, storage, and use.

FAQ

Frequently asked questions about Cybersecurity AI Dataset.

How is this different from general AI training data?

Cybersecurity AI data isn't just text or images — the relationships and context between threats matter, including malware, IoCs, attack groups, attack techniques, and vulnerabilities. SANDS Lab structures feature information and labels that security AI can learn from, based on real threat analysis technology.

What cybersecurity domains have you built data for?

Through public-sector projects from 2021 to 2024, we've built datasets across 8 domains: malware, incident response, application security, proactive security monitoring, threat profiling, threat hunting, threat intelligence, and recent incidents.

How do you validate data quality?

We inspect data quality at the collection, processing, and labeling stages, going through error analysis and improvement procedures. We then validate the data's effectiveness by training a real AI model on the built dataset and checking performance metrics such as precision, recall, and F1 score.

Do you support standard formats like STIX?

In our public-sector projects, we structured threat data using structured formats such as JSON and STIX 2.1, and built a data structure designed with interoperability with external security systems in mind.

Can you build a new dataset for a specific purpose?

SANDS Lab has experience building the entire data pipeline — from raw threat data collection through analysis, metadata generation, labeling, quality validation, and AI model validation. The actual build scope is designed based on the research purpose, data rights, and the required threat domain.

Can I purchase the dataset described on this page directly?

This page is an R&D page introducing SANDS Lab's cybersecurity AI data building and validation capability, rather than a specific commercial dataset product. The scope of data provision, joint research, or build collaboration is discussed separately based on data rights and project terms.

Let’s design the data your security AI needs —together.

From cyber threat data collection, processing, and labeling to quality validation and AI model demonstration — put SANDS Lab's data and AI research capability to work.

How does this product behave in your environment?

Whether you're exploring, evaluating, or rolling out, you connect directly with a SANDS Lab solutions engineer. Clear every question before contract — that's the point.

FOR EVALUATORS

Product evaluation & PoC

Real-data PoCs, technical deep-dive sessions, and custom integration scoping. Everything you'd need to validate technical fit before the paperwork starts.

FOR RESEARCHERS

Technical collaboration & licensing

If you want to use the product in an academic benchmark or co-authored paper, we support research licenses and the underlying datasets. Co-authorship is on the table.