Official PyTorch implementation of:
ESICA: A Scalable Framework for Text-Guided 3D Medical Image Segmentation
Text-guided 3D medical image segmentation offers a flexible alternative to class-based and spatial prompt-based models by allowing users to specify regions of interest directly in natural language. This paradigm avoids reliance on predefined label sets, reduces ambiguous outputs, and aligns more naturally with clinical workflows. However, existing text-guided frameworks are often computationally expensive, exhibit weak text–volume feature alignment, and fail to capture fine anatomical details. We propose ESICA, a lightweight and scalable framework that addresses these challenges through three innovations: (1) a similarity-matrix-based mask prediction formulation that enhances semantic alignment, (2) an efficient decomposed decoder with adapter modules for accurate volumetric decoding, and (3) a two-pass refinement strategy that sharpens boundaries and resolves uncertain regions. To improve training stability and generalization, ESICA adopts a two-stage scheme consisting of positive-only pretraining followed by balanced fine-tuning. On the CVPR-BiomedSegFM benchmark spanning five imaging modalities (CT, MRI, PET, ultrasound, and microscopy), ESICA achieves state-of-the-art segmentation accuracy, while the compact ESICA4-Lite variant attains similar segmentation performance with substantially fewer parameters, yielding a superior efficiency–accuracy trade-off. Our framework advances text-guided segmentation toward efficient, scalable, and clinically deployable systems.
- Python==3.12.11
- torch==2.8.0
- torchvision==0.23.0
- monai==1.5.0
- deepspeed==0.17.4
First, clone the repository to your local machine:
git clone https://github.com/mirthAI/ESICA.git
cd ESICATo install the required packages, you can use the following command:
conda create -n ESICA python=3.12.11
conda activate ESICA
pip install -r requirements.txtIn the paper, we train and evaluate our model on CVPR-BiomedSegFM datasets.
Use the following command to prepare the data:
sh scripts/prepare_data.shThe pre-trained model weights are available on Hugging Face: ESICA Pre-trained Weights.
To train the model, use the following command:
sh scripts/train.shor you can train the lightweight version of the model using the following command:
sh scripts/train_lite.shTo evaluate the model, use the following command:
sh scripts/eval.shThe code is only for research purposes. If you have any questions regarding how to use this code, feel free to contact Yu Xin at yu.xin@ufl.edu.
Kindly cite the following papers if you use our code.
