HbHAIA - Dataset Zoo

Dataset 1

For each of the compression rate, the underlying plaintext dataset is different.

  • Version 1. Cybersecurity data. The compression rate is delta = 3. Only the HbHAI protected version is provided Dataset1.zip (one single ZIP file containing training and validation data in ZIP and TGZ formats). The training set contains 4000 objects (2000 per class). The validation set contains objects (200 per class). Each object is defined by 49 955 features.
  • Version 2. Cybersecurity data. The compression rate is delta = 10. Only the HbHAI protected version is provided. Dataset1-delta10.tgz (one single TGZ file containing training and validation data). The training set contains 2000 objects (1000 per class). The validation set contains 200 objects (100 per class). Each object is defined by 49 955 features.
  • Version 3. Cybersecurity data. The compression rate is delta = 10. The plaintext version of this data set is provided on Dataset1-plain.tgz (one single TGZ file containing training and validation data). The training set contains 4000 objects (2000 per class). The validation set contains objects (200 per class). Each object is defined by 48 830 features. The corresponding HbHAI-protected version is available on a Dataset1-hbhai-delta10.tgz (one single TGZ file containing training and validation data)

Dataset 2

We consider here the reference dataset for deep learning evaluation, especially with neural networks. Zalando Fashion MNIST proposed by Zalando Research here. All HbHAI-protected versions keep the original structure as detailed here

  • Compression rate Delta = 3. Delta-3.zip (one single ZIP file containing training and validation data in GZ format).
  • Compression rate Delta = 6. Delta-6.zip (one single ZIP file containing training and validation data in GZ format).
  • Compression rate Delta = 6 for three different 256-bit keys: delta6-key1.tgz, delta6-key2.tgz, delta6-key3.tgz.

Third-party Datasets

We hope that contributors will propose their own datasets to test HbHAI techniques. We favour especially massive datasets which enable to harness the full power of HbHAI techniques.