Posted on Leave a comment

Creating EMBeD, the Embedded Malware Benchmark Dataset

Hello there! My name is Dávid Maliga, a first-year PhD student – under the supervision of Prof. Levente Buttyán – in the CrySyS lab, and I would like to share with you our latest work, which we just recently presented at this year’s CITDS conference (2026 IEEE 4th Conference on Information Technology and Data Science) in the city of Debrecen, at the beautiful campus of the University of Debrecen.

Repeatability of expreiments is an important requirement for any scientific approach. Repeatability means that if someone wants to apply an existing approach to their own problem, they should be able to do so reliably, by following the documented methodology. The same methodology under which the original approach was developed and validated. This means, to verify an approach, repeating the same measurements on the same dataset should produce exactly the same results.

Unfortunately, in the field of IoT malware detection, the current tendency is that individual research groups use their own, proprietary datasets. This renders reproducibility difficult, because even if we succesfully reproduce their approach, we don’t have their data to validate it. If we consider machine-learning based approaches, even if two approaches rely on the same dataset, the specific train–test split could be unavailable. This way we couldn’t tell whether the difference in results come from the approaches or the split. This raises the question of whether results obtained in different settings are genuinely comparable. Without a shared reference point, it becomes hard to say whether one approach outperforms the other.

To address this issue, we have developed a methodology for creating EMBeD (Embedded Malware Benchmark Dataset), a publicly available benchmark dataset specifically for IoT malware binaries. This method processes existing malware datasets to extract metadata and filter out unreliable or underrepresented samples. It also verifies malware family labels and balances the dataset in terms of number of samples per malware family.

Workflow

Data preparation

The first step is to build a malware metadata database for the next stages of the workflow. Malware samples from existing public datasets are processed through a pipeline that extracts basic information about each file, such as its hash, architecture, bitness and entropy, as well as various statistics from VirusTotal scan results. Samples that appear to be benign or have underrepresented properties are filtered out. The result is a clean, detailed database of potential malware samples that can be added to the benchmark dataset.

Enhancing family labeling

Next, a reliable family label is required for each sample. This information can be extracted from the VirusTotal scan results by aggregating the votes and labels provided by various antivirus engines. As the antivirus engines often disagree about the family label, the votes are weighted according to the historical accuracy of the engines to provide an initial family label for each sample.

These initial labels are then validated using a similarity graph. In this structure, each node represents a sample, and two nodes are connected if, and only if their corresponding samples are similar. We consider two samples similar if their TLSH difference score is below a certain threshold, where TLSH is well-known fuzzy hashing scheme that can quantify the dissimilarity of two samples. For each sample, the weighted initial family labels of its neighbours in the similarity graph are aggregated, with those that have lower dissimilarity score getting larger weight. If the new aggregated label differs from the sample’s initially assigned label, then we consider the initial label to be inconsistent, and the sample is removed. This process is repeated until all remaining labels are consistent.

Balancing the family sizes

Finally, the dataset needs to be balanced, as some malware families have far more samples than others. For large families, we select a representative subset of the samples conatining a target number of samples. For small families below the target size, we search for additional samples on VirusTotal using the LiveHunt and RetroHunt features. We also perform similarity search in Ukatemi’s Kaibou Repo, a large repository of around 800 million malware binaries with a fast similiraty search engine. This process is repeated until family sizes become balanced.

Preliminary results

To verify that our method was working, we created a proof-of-concept (PoC) benchmark dataset, aiming to collect 100 samples per family. To achieve this, we applied the process to 67,800 MIPS-based IoT malware samples. After labelling and balancing, we achieved an evenly sized benchmark dataset covering seven malware families: mirai, gafgyt, hajime, kaiji, tsunami, ddostf and dofloo. With this small-scale PoC dataset, we demonstrated that the entire workflow is effective and can produce a benchmark dataset.

Result

Our goal for the future is to expand this project and publish a much larger version of EMBeD as soon as possible. This will involve extending the scope beyond MIPS to include other architectures commonly used in IoT devices, such as ARM. This would make the benchmark useful for a much wider range of malware detection research across the IoT domain.

If you are interested in the methodology and would like to know more details, the paper presented at the conference is available here. It should be cited as follows: Dávid Maliga, Dorottya Papp and Levente Buttyán, Creating EMBeD, the Embedded Malware Benchmark Dataset, IEEE Conference on Information Technology and Data Science, 2026.

We would like to thank VirusTotal for providing an academic license, and Ukatemi Technologies for granting access to their Kaibou malware repository.

Posted on Leave a comment

CAN We Trust Your Results? A Cross-Dataset Study of Automotive IDS Evaluation

Our new paper is available now here.
You can cite it like this: B. Koltai, G. Ács and A. Gazdag, CAN We Trust Your Results? A Cross-Dataset Study of Automotive IDS Evaluation, Euro S&P – ACSW26, 2026.

Modern vehicles are no longer just mechanical machines. They are complex computer systems on wheels, with many electronic control units constantly talking to each other. One of the main communication systems inside vehicles is the Controller Area Network, or CAN bus.
The CAN bus was designed to be fast and reliable, but not necessarily secure. As cars have become more connected, this has become a serious concern. If an attacker gains access to the vehicle network, they may try to inject fake messages, block legitimate ones, or modify vehicle signals. This is why many researchers have worked on intrusion detection systems, or IDSs, for automotive networks.
These systems are meant to detect suspicious activity on the CAN bus. But there is an important question behind this work:

If an IDS performs well in one study, can we trust that it will also work well in another vehicle, another dataset, or another attack scenario? Or in other words: Are some current methods too dependent on the specific datasets they used?

Continue reading CAN We Trust Your Results? A Cross-Dataset Study of Automotive IDS Evaluation

Posted on Leave a comment

GAME: Genetic Algorithm for Malware Evasion

This post summarizes the work we presented at the 15th Conference of PhD Students in Computer Science (CSCS), 2026 in Szeged. My name is Jószef Sándor, I’m a PhD student in the CrySyS Lab, mentored by Prof. Levente Buttyán. A BSc student, Bence Kovács, was also involved in this project, mostly working on the implementation of our ideas.

Imagine a lightweight malware detector running on a small IoT device. It’s fast, compact, and accurate on test data. Now imagine an attacker who tweaks a malware file just enough to fool the detector – while the malware still works exactly as intended. This is a rational move for the attacker: building a whole new malware is difficult and expensive, so reusing existing malware – even at the binary level – while evading detection is a much cheaper path to success.

Continue reading GAME: Genetic Algorithm for Malware Evasion

Posted on Leave a comment

Increasing storage efficiency in large malware repositories by differential storage and incremental clustering of samples

Hi there! I’m Dávid Maliga, a PhD student in the CrySyS lab, under the supervision of Prof. Levente Buttyán, and I have recently presented our extended abstract at this year’s CSCS conference (15th Conference of PhD Students in Computer Science) in the wonderful city of Szeged. This work was part of an R&D project that focused on designing efficient differential storage schemes for large malware repositories (e.g., VirusTotal, MalwareBazaar, or Kaibou Repo). Our project partner was Ukatemi Technologies, and the project was funded by the National Research, Development and Innovation Office of Hungary under grant number 2023-1.1.1PIACI_FÓKUSZ-2024-00030.

The main idea of our work stems from the observation that many malware samples are not entirely new, but rather minor modifications of existing members of a malware family. Thus, if we organize similar samples into groups (clusters), it suffices to store one “complete” sample per group – a representative sample for the whole cluster – with the rest stored as differences with respect to the completely stored sample; this is known as differential storage.

Continue reading Increasing storage efficiency in large malware repositories by differential storage and incremental clustering of samples

Posted on Leave a comment

Locked Shields Partner Run 2026 – kiberbiztonság az egyetemi gyakorlatban

Az első 3 között végzett a magyar csapat a Locked Shields 2026 Partner Runon! (Pontosabb eredményt nem hírdettek a szervezők.)

A Locked Shields minden évben a világ legnagyobb és egyik legösszetettebb valós idejű kibervédelmi gyakorlataként teszi próbára a NATO szövetségesek és partnerek kiberbiztonsági szakértőit. Ennek az gyakorlatnak a “főpróbája” a Locked Shields Partner Run, amelyre a hazai egyetemek több mint tíz éve delegálnak oktatókat és hallgatókat. Nagyon jó példája ez az egyetemi kooperációnak, amelynek oktatás-kutatási és innovációs eredményeit ma már az oktatásban is használjuk. A Partner Run egyben lehetőséget terem a hallgatóknak arra, hogy az elméleti tudásukat valós körülmények között próbára tegyék. A gyakorlatot a NATO Kooperatív Kibervédelmi Kiválósági Központ (CCDCOE) szervezi Tallinnból, Észtországból.

A gyakorlat alapvetően egy “Red Team vs. Blue Team” struktúrára épül: a vörös csapat szerepe a támadás szimulálása, míg a kék csapatok – nemzeti vagy nemzetközi formációk – célja a védelmi stratégiák és technikai megoldások összehangolt alkalmazása a támadások kivédésére, és a támadó beazonosítására. A gyakorlat során a résztvevők nem csupán műszaki kihívásokkal néznek szembe, hanem a stratégiai döntéshozatal, jogi kérdések, igazságügyi feladatok és a krízis kommunikáció területén is komplex, dinamikusan változó helyzeteket kell kezelniük.

A gyakorlaton való közös részvétel már egyfajta tavaszi hagyomány a Budapesti Műszaki és Gazdaságtudományi Egyetem, az Óbudai Egyetem és a Nemzeti Közszolgálati Egyetem hallgatói és oktatói számára. Ehhez tavaly csatlakoztak az Eötvös Loránd Tudományegyetem, majd idén a Pécsi Tudományegyetem, illetve a győri Széchényi István Egyetem hallgatói is.

A szervezők által közzétett hír elérhető itt.

Posted on Leave a comment

CrySyS dataset of CAN traffic logs containing fabrication and masquerade attacks

Our paper introducing a new CAN dataset is now available in Nature: Scientific Data.
The dataset contains 26 recordings of benign network traffic, amounting to more than 2.5 hours of traffic. We performed two attacks (injection and modification) with different configurations multiple times on each benign trace to create a comprehensive set of traffic logs. The dataset structure was explicitly designed with machine learning applications in mind.

Continue reading CrySyS dataset of CAN traffic logs containing fabrication and masquerade attacks

Posted on Leave a comment

Gépi Tanulás & Adatbiztonsági Védekezések

Ez a blogposzt a második egy két részes sorozatból mely a gépi tanulás adatbiztonsági kockázatairól kíván közérthető nyelven egy átfogó képet nyújtani. Ez a bejegyzés a létező védekezéseket tárgyalja, míg az előző a lehetséges támadásokat mutatta be.

Continue reading Gépi Tanulás & Adatbiztonsági Védekezések

Posted on Leave a comment

Gépi Tanulás & Adatbiztonsági Támadások

Ez a blogposzt az első egy két részes sorozatból mely a gépi tanulás adatbiztonsági kockázatairól kíván közérthető nyelven egy átfogó képet nyújtani. Ez a bejegyzés néhány létező támadást fed le, míg a következő a lehetséges védekezéseket mutatja be.

Continue reading Gépi Tanulás & Adatbiztonsági Támadások

Posted on Leave a comment

Differenciális Adatvédelem

Napjainkban az információs technológiák kiemelkedően fontos szerepet töltenek be mindannyiunk életében. Mind a munkánk, mind a magánéletünk során folyamatos érintkezésben vagyunk digitális szolgáltatásokkal, kezdve a pénzügyeinktől a társkeresésen át a magánbeszélgetéseinkig. Ezek és hasonló szolgáltatások használata során ugyanakkor digitális lábnyomokat hagyunk magunk után, amik komoly adatvédelmi kockázatot is jelenthetnek. A bizalmasan megosztott véleményünk ugyanolyan érzékeny adat, mint a pénzügyi helyzetünk és a szexuális hovatartozásunk, emiatt elengedhetetlen az adataink megfelelő védelme.

Európában a személyes jellegű adatokat törvények védik, amik például azok explicit megosztását is korlátozzák. Ilyen az Általános adatvédelmi rendelet (GDPR), vagy a Digitális szolgáltatások jogszabály (DSA) is. Ennek ellenére rendszeresen történnek adatszivárgások, amiket külső támadók és belső hibák egyaránt okoznak. Látható tehát, hogy pusztán a jogi védelem nem elégséges, és kiegészítő megoldások használata nélkülözhetetlen. Ilyen például a PET (Privacy Enhancing Technologies), ami olyan eljárásokat foglal magában, melyek célja, hogy megvédjék az adatokat a jogosulatlan hozzáféréstől.

Ebben a cikkben a gépi tanulás adatvédelmi kockázataira fókuszálunk, és számos támadás ismertetése mellett bemutatjuk az egyik legelterjedtebb és leghatásosabb védekezési PET mechanizmust, az úgynevezett differenciális adatvédelmet.

Continue reading Differenciális Adatvédelem

Posted on Leave a comment

Post-Quantum Cryptography Standardization: A New Milestone

milestone Post-Quantum Cryptography Standardization: A New Milestone

In some of our previous posts, we have already touched upon why the development of quantum computers poses challenges to the field of information security and how the standardisation bodies, most notably NIST, prepare for the post-quantum era of computing. This process reached its next milestone yesterday when NIST has announced which key-establishment mechanism and digital signature schemes will be standardized soon.

Continue reading Post-Quantum Cryptography Standardization: A New Milestone