Compare the original MRI, ground truth, and SAM2-Large predictions before and after fine-tuning.
Four views, one slice
Left to right: MRI slice along the z-axis, ground-truth tumor mask, prediction without fine-tuning, and prediction with fine-tuning.
Approach
A 3D problem. A 2D foundation model.
BraTS NIfTI volumes are converted into 2D slices. FLAIR inputs are replicated across three channels, then treated as a pseudo-video for continuous segmentation.
Processing pipeline
Prepare the volumes
Read multimodal BraTS MRI volumes and slice them along the z-axis.
Construct model inputs
Replicate FLAIR into three channels and save preprocessed .npz files.
Fine-tune with LoRA
Apply rank-16 adapters to selected attention and MLP layers.
Propagate across slices
Use forward and backward pseudo-video propagation from an initial labeled slice.
Method highlights
Pre-slicing improves training efficiency and GPU utilization.
FLAIR preserves tumor-sensitive information while matching the model’s input format.
Only about 0.3% of model parameters are fine-tuned, according to the project write-up.
My contributions
Designed the 3D-to-2D preprocessing pipeline.
Implemented parameter-efficient LoRA fine-tuning.
Built the inference and pseudo-video propagation pipeline.
Findings & limitations
Fine-tuning makes the difference.
Observed results
Lower-resolution training reduced training time. Fine-tuning remained necessary for acceptable segmentation quality; direct prediction without adaptation produced unsatisfactory results.
What remains open
The model targets brain tumors and still needs one manually labeled slice as a prompt. Future work could explore higher-resolution inputs and a larger trainable parameter budget.
The project report is being finalized and will be added when ready.