Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
94 changes: 44 additions & 50 deletions Tutorials/DeepDives/ConfigTutorial.ipynb

Large diffs are not rendered by default.

22 changes: 22 additions & 0 deletions Tutorials/DeepDives/EvaluateTutorial.ipynb
Original file line number Diff line number Diff line change
@@ -1,5 +1,27 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "c0a1e6e7",
"metadata": {},
"source": [
"> **Note:** This environment variable is required for fully deterministic CuBLAS ops on CUDA >= 10.2 when `reproducible=True` is set below. Without it, PyTorch raises a `RuntimeError` instead of training deterministically. It must be set before `torch` is imported. See the [README FAQ](../../README.md#reproducibility-and-cublas_workspace_config) for details."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "679c00e4",
"metadata": {
"vscode": {
"languageId": "shellscript"
}
},
"outputs": [],
"source": [
"%env CUBLAS_WORKSPACE_CONFIG=:16:8"
]
},
{
"cell_type": "markdown",
"id": "89f4ba9f",
Expand Down
95 changes: 54 additions & 41 deletions Tutorials/DeepDives/ExplainStep.ipynb
Original file line number Diff line number Diff line change
@@ -1,5 +1,36 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "1bc0ea1c",
"metadata": {},
"source": [
"> **Note:** This environment variable is required for fully deterministic CuBLAS ops on CUDA >= 10.2 when `reproducible=True` is set below. Without it, PyTorch raises a `RuntimeError` instead of training deterministically. It must be set before `torch` is imported. See the [README FAQ](../../README.md#reproducibility-and-cublas_workspace_config) for details."
]
},
{
"cell_type": "code",
"execution_count": 1,
"id": "c29ff66a",
"metadata": {
"vscode": {
"languageId": "shellscript"
}
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"env: CUBLAS_WORKSPACE_CONFIG=:16:8\n"
]
}
],
"source": [
"# Setting CUBLAS for reproducibility if you train on GPU\n",
"%env CUBLAS_WORKSPACE_CONFIG=:16:8"
]
},
{
"cell_type": "markdown",
"id": "74aebb64",
Expand All @@ -20,13 +51,7 @@
"### 2) How To Perform the Explain Step\n",
"First, we need to run an AUTOENCODIX pipeline, `Ontix` is a very good choice here.\n",
"#### ❗❗ Requirements: Getting Tutorial Data ❗❗\n",
"To follow along, please download the date from the link below (1GB):\n",
"\n",
"https://cloud.scadsai.uni-leipzig.de/index.php/s/HcaB8csTctGimQQ/download/OntixTutorialData.zip\n",
"\n",
"After downloading:\n",
"- from the root of the repository, create the folders `data/raw` if not created yet\n",
"- move the donwloaded files there\n",
"The data for this tutorial is hosted on Hugging Face Hub ([autoencodix/tcga](https://huggingface.co/datasets/autoencodix/tcga)) and is downloaded automatically in the cell below on first run.\n",
"\n",
"#### Extra 2: Get correct path\n",
"We assume you are in the root of the package. The following code ensures that the correct paths are used.\n",
Expand All @@ -41,41 +66,10 @@
},
{
"cell_type": "code",
"execution_count": 1,
"id": "c29ff66a",
"metadata": {
"vscode": {
"languageId": "shellscript"
}
},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"env: CUBLAS_WORKSPACE_CONFIG=:16:8\n"
]
}
],
"source": [
"# Setting CUBLAS for reproducibility if you train on GPU\n",
"%env CUBLAS_WORKSPACE_CONFIG=:16:8"
]
},
{
"cell_type": "code",
"execution_count": 2,
"execution_count": null,
"id": "2f748e0a",
"metadata": {},
"outputs": [
{
"name": "stdout",
"output_type": "stream",
"text": [
"Changed to: /home/ewald/Github/acx_main_releases/autoencodix_package\n"
]
}
],
"outputs": [],
"source": [
"import os\n",
"\n",
Expand All @@ -84,7 +78,26 @@
"if d not in p:\n",
" raise FileNotFoundError(f\"'{d}' not found in path: {p}\")\n",
"os.chdir(os.sep.join(p.split(os.sep)[: p.split(os.sep).index(d) + 1]))\n",
"print(f\"Changed to: {os.getcwd()}\")\n"
"print(f\"Changed to: {os.getcwd()}\")\n",
"\n",
"# ---------------------------------------------------------------------\n",
"# Data is hosted on Hugging Face Hub and downloaded (and locally cached)\n",
"# automatically, then placed under the paths used below\n",
"# ---------------------------------------------------------------------\n",
"import shutil\n",
"from huggingface_hub import hf_hub_download\n",
"\n",
"HF_REPO_ID = \"autoencodix/tcga\"\n",
"os.makedirs(\"data/raw\", exist_ok=True)\n",
"for hf_filename, local_name in [\n",
" (\"rna.parquet\", \"combined_rnaseq_formatted.parquet\"),\n",
" (\"methylation.parquet\", \"combined_meth_formatted.parquet\"),\n",
" (\"clinical.parquet\", \"combined_clin_formatted.parquet\"),\n",
"]:\n",
" downloaded_path = hf_hub_download(\n",
" repo_id=HF_REPO_ID, repo_type=\"dataset\", filename=hf_filename\n",
" )\n",
" shutil.copyfile(downloaded_path, os.path.join(\"data/raw\", local_name))\n"
]
},
{
Expand Down
44 changes: 33 additions & 11 deletions Tutorials/DeepDives/HyperparameterOptimizationOptunaTutorial.ipynb
Original file line number Diff line number Diff line change
@@ -1,5 +1,27 @@
{
"cells": [
{
"cell_type": "markdown",
"id": "95995f18",
"metadata": {},
"source": [
"> **Note:** This environment variable is required for fully deterministic CuBLAS ops on CUDA >= 10.2 when `reproducible=True` is set below. Without it, PyTorch raises a `RuntimeError` instead of training deterministically. It must be set before `torch` is imported. See the [README FAQ](../../README.md#reproducibility-and-cublas_workspace_config) for details."
]
},
{
"cell_type": "code",
"execution_count": null,
"id": "f54813b1",
"metadata": {
"vscode": {
"languageId": "shellscript"
}
},
"outputs": [],
"source": [
"%env CUBLAS_WORKSPACE_CONFIG=:16:8"
]
},
{
"metadata": {},
"cell_type": "markdown",
Expand Down Expand Up @@ -278,7 +300,7 @@
"name": "stderr",
"output_type": "stream",
"text": [
"\u001B[32m[I 2026-06-04 14:22:48,108]\u001B[0m A new study created in memory with name: autoencodix-optimization\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:48,108]\u001b[0m A new study created in memory with name: autoencodix-optimization\u001b[0m\n"
]
},
{
Expand All @@ -292,7 +314,7 @@
"name": "stderr",
"output_type": "stream",
"text": [
"\u001B[32m[I 2026-06-04 14:22:48,746]\u001B[0m Trial 0 finished with value: 208.5494140625 and parameters: {'batch_size': 1783, 'drop_p': 0.6482920440979423, 'enc_factor': 1, 'weight_decay': 0.0001619311091244073, 'beta': 7.595132328682394e-05, 'learning_rate': 2.3407464805767515e-05, 'n_layers': 2}. Best is trial 0 with value: 208.5494140625.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:48,746]\u001b[0m Trial 0 finished with value: 208.5494140625 and parameters: {'batch_size': 1783, 'drop_p': 0.6482920440979423, 'enc_factor': 1, 'weight_decay': 0.0001619311091244073, 'beta': 7.595132328682394e-05, 'learning_rate': 2.3407464805767515e-05, 'n_layers': 2}. Best is trial 0 with value: 208.5494140625.\u001b[0m\n"
]
},
{
Expand All @@ -311,7 +333,7 @@
"text": [
"/Users/lucathale-bombien/PycharmProjects/autoencodix_package/.venv/lib/python3.12/site-packages/sklearn/linear_model/_sag.py:348: ConvergenceWarning: The max_iter was reached which means the coef_ did not converge\n",
" warnings.warn(\n",
"\u001B[32m[I 2026-06-04 14:22:49,627]\u001B[0m Trial 1 finished with value: 204.61236328125 and parameters: {'batch_size': 1499, 'drop_p': 0.35709072680760295, 'enc_factor': 3, 'weight_decay': 0.00047509237210306113, 'beta': 0.12921621108432518, 'learning_rate': 6.573686655138327e-05, 'n_layers': 5}. Best is trial 1 with value: 204.61236328125.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:49,627]\u001b[0m Trial 1 finished with value: 204.61236328125 and parameters: {'batch_size': 1499, 'drop_p': 0.35709072680760295, 'enc_factor': 3, 'weight_decay': 0.00047509237210306113, 'beta': 0.12921621108432518, 'learning_rate': 6.573686655138327e-05, 'n_layers': 5}. Best is trial 1 with value: 204.61236328125.\u001b[0m\n"
]
},
{
Expand All @@ -330,7 +352,7 @@
"text": [
"/Users/lucathale-bombien/PycharmProjects/autoencodix_package/.venv/lib/python3.12/site-packages/sklearn/linear_model/_sag.py:348: ConvergenceWarning: The max_iter was reached which means the coef_ did not converge\n",
" warnings.warn(\n",
"\u001B[32m[I 2026-06-04 14:22:50,426]\u001B[0m Trial 2 finished with value: 203.7865625 and parameters: {'batch_size': 236, 'drop_p': 0.603420759160562, 'enc_factor': 3, 'weight_decay': 0.0017169565852473863, 'beta': 6.955392321661599e-05, 'learning_rate': 6.200203677164716e-05, 'n_layers': 5}. Best is trial 2 with value: 203.7865625.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:50,426]\u001b[0m Trial 2 finished with value: 203.7865625 and parameters: {'batch_size': 236, 'drop_p': 0.603420759160562, 'enc_factor': 3, 'weight_decay': 0.0017169565852473863, 'beta': 6.955392321661599e-05, 'learning_rate': 6.200203677164716e-05, 'n_layers': 5}. Best is trial 2 with value: 203.7865625.\u001b[0m\n"
]
},
{
Expand All @@ -347,7 +369,7 @@
"name": "stderr",
"output_type": "stream",
"text": [
"\u001B[32m[I 2026-06-04 14:22:50,808]\u001B[0m Trial 3 finished with value: 210.5692578125 and parameters: {'batch_size': 3971, 'drop_p': 0.28208176034331855, 'enc_factor': 4, 'weight_decay': 0.03202997570029044, 'beta': 2.3315244870014054, 'learning_rate': 2.188652665746547e-05, 'n_layers': 2}. Best is trial 2 with value: 203.7865625.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:50,808]\u001b[0m Trial 3 finished with value: 210.5692578125 and parameters: {'batch_size': 3971, 'drop_p': 0.28208176034331855, 'enc_factor': 4, 'weight_decay': 0.03202997570029044, 'beta': 2.3315244870014054, 'learning_rate': 2.188652665746547e-05, 'n_layers': 2}. Best is trial 2 with value: 203.7865625.\u001b[0m\n"
]
},
{
Expand All @@ -364,7 +386,7 @@
"name": "stderr",
"output_type": "stream",
"text": [
"\u001B[32m[I 2026-06-04 14:22:51,727]\u001B[0m Trial 4 finished with value: 202.02994140625 and parameters: {'batch_size': 802, 'drop_p': 0.7903282530864718, 'enc_factor': 1, 'weight_decay': 0.0004835378777285505, 'beta': 5.58903952498834, 'learning_rate': 0.0013572540301313666, 'n_layers': 4}. Best is trial 4 with value: 202.02994140625.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:51,727]\u001b[0m Trial 4 finished with value: 202.02994140625 and parameters: {'batch_size': 802, 'drop_p': 0.7903282530864718, 'enc_factor': 1, 'weight_decay': 0.0004835378777285505, 'beta': 5.58903952498834, 'learning_rate': 0.0013572540301313666, 'n_layers': 4}. Best is trial 4 with value: 202.02994140625.\u001b[0m\n"
]
},
{
Expand All @@ -383,7 +405,7 @@
"text": [
"/Users/lucathale-bombien/PycharmProjects/autoencodix_package/.venv/lib/python3.12/site-packages/sklearn/linear_model/_sag.py:348: ConvergenceWarning: The max_iter was reached which means the coef_ did not converge\n",
" warnings.warn(\n",
"\u001B[32m[I 2026-06-04 14:22:52,225]\u001B[0m Trial 5 finished with value: 202.1137109375 and parameters: {'batch_size': 1380, 'drop_p': 0.6178508349134253, 'enc_factor': 5, 'weight_decay': 1.183458707441061e-05, 'beta': 0.3168588850299515, 'learning_rate': 0.09024940671927402, 'n_layers': 4}. Best is trial 4 with value: 202.02994140625.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:52,225]\u001b[0m Trial 5 finished with value: 202.1137109375 and parameters: {'batch_size': 1380, 'drop_p': 0.6178508349134253, 'enc_factor': 5, 'weight_decay': 1.183458707441061e-05, 'beta': 0.3168588850299515, 'learning_rate': 0.09024940671927402, 'n_layers': 4}. Best is trial 4 with value: 202.02994140625.\u001b[0m\n"
]
},
{
Expand All @@ -400,7 +422,7 @@
"name": "stderr",
"output_type": "stream",
"text": [
"\u001B[32m[I 2026-06-04 14:22:52,943]\u001B[0m Trial 6 finished with value: 205.6899609375 and parameters: {'batch_size': 1241, 'drop_p': 0.7103513956063396, 'enc_factor': 1, 'weight_decay': 0.0006188339116517525, 'beta': 2.8286096492960997, 'learning_rate': 0.0001494364677378836, 'n_layers': 3}. Best is trial 4 with value: 202.02994140625.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:52,943]\u001b[0m Trial 6 finished with value: 205.6899609375 and parameters: {'batch_size': 1241, 'drop_p': 0.7103513956063396, 'enc_factor': 1, 'weight_decay': 0.0006188339116517525, 'beta': 2.8286096492960997, 'learning_rate': 0.0001494364677378836, 'n_layers': 3}. Best is trial 4 with value: 202.02994140625.\u001b[0m\n"
]
},
{
Expand All @@ -417,7 +439,7 @@
"name": "stderr",
"output_type": "stream",
"text": [
"\u001B[32m[I 2026-06-04 14:22:53,322]\u001B[0m Trial 7 finished with value: 118.65896484375 and parameters: {'batch_size': 644, 'drop_p': 0.017430262083267367, 'enc_factor': 4, 'weight_decay': 7.022834985764343e-05, 'beta': 0.0003919944843607491, 'learning_rate': 0.0009253214657642581, 'n_layers': 2}. Best is trial 7 with value: 118.65896484375.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:53,322]\u001b[0m Trial 7 finished with value: 118.65896484375 and parameters: {'batch_size': 644, 'drop_p': 0.017430262083267367, 'enc_factor': 4, 'weight_decay': 7.022834985764343e-05, 'beta': 0.0003919944843607491, 'learning_rate': 0.0009253214657642581, 'n_layers': 2}. Best is trial 7 with value: 118.65896484375.\u001b[0m\n"
]
},
{
Expand All @@ -436,7 +458,7 @@
"text": [
"/Users/lucathale-bombien/PycharmProjects/autoencodix_package/.venv/lib/python3.12/site-packages/sklearn/linear_model/_sag.py:348: ConvergenceWarning: The max_iter was reached which means the coef_ did not converge\n",
" warnings.warn(\n",
"\u001B[32m[I 2026-06-04 14:22:53,800]\u001B[0m Trial 8 finished with value: 190.7475 and parameters: {'batch_size': 2406, 'drop_p': 0.13205571741522915, 'enc_factor': 3, 'weight_decay': 0.00629554655842228, 'beta': 4.111559438268768e-05, 'learning_rate': 0.0004531311844091796, 'n_layers': 4}. Best is trial 7 with value: 118.65896484375.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:53,800]\u001b[0m Trial 8 finished with value: 190.7475 and parameters: {'batch_size': 2406, 'drop_p': 0.13205571741522915, 'enc_factor': 3, 'weight_decay': 0.00629554655842228, 'beta': 4.111559438268768e-05, 'learning_rate': 0.0004531311844091796, 'n_layers': 4}. Best is trial 7 with value: 118.65896484375.\u001b[0m\n"
]
},
{
Expand All @@ -455,7 +477,7 @@
"text": [
"/Users/lucathale-bombien/PycharmProjects/autoencodix_package/.venv/lib/python3.12/site-packages/sklearn/linear_model/_sag.py:348: ConvergenceWarning: The max_iter was reached which means the coef_ did not converge\n",
" warnings.warn(\n",
"\u001B[32m[I 2026-06-04 14:22:54,334]\u001B[0m Trial 9 finished with value: 136.542578125 and parameters: {'batch_size': 1771, 'drop_p': 0.04495811305147845, 'enc_factor': 3, 'weight_decay': 0.0045204178469246655, 'beta': 0.012283854745700366, 'learning_rate': 0.060031476336384136, 'n_layers': 4}. Best is trial 7 with value: 118.65896484375.\u001B[0m\n"
"\u001b[32m[I 2026-06-04 14:22:54,334]\u001b[0m Trial 9 finished with value: 136.542578125 and parameters: {'batch_size': 1771, 'drop_p': 0.04495811305147845, 'enc_factor': 3, 'weight_decay': 0.0045204178469246655, 'beta': 0.012283854745700366, 'learning_rate': 0.060031476336384136, 'n_layers': 4}. Best is trial 7 with value: 118.65896484375.\u001b[0m\n"
]
},
{
Expand Down
Loading
Loading