|
Hello, thanks for the new tutorials. This is exactly what I was looking for! It makes the whole process of fine-tuning much easier to understand. So with in the default notebook (I haven't changed anything): So I tried training with more (15) epochs and not freeze the encoder, which also did not produce any more checkpoint files. Then I included 'save_top_k=5, save_last=True' into the callback function, which produces multiple 'best-epoch=XX.ckpt' files. I used a newer .ckpt file, i.e. 'best-epoch=04.ckpt', but it also only predicts 'no burned'. Something is going wrong, especially as the Loss is NaN. These are my metrics for the third test: [{'test/loss': nan, I'm training on NVIDIA RTXA6000-48Q with CUDA Version: 12.2 Best regards, |
Replies: 6 comments 2 replies
|
Hi, @Atmoboran. Just to inform you, I'm checking and trying to reproduce your issue. |
|
here I have my ipynb file attached. So if that helps you can look at my outputs. prithvi_v2_eo_300_tl_unet_burnscars.zip Thank you! |
|
@Atmoboran The notebook is failing for me, but for others reasons. I'll run it with a different environment: |
|
So I have found the problem while watching @paolo-fraccaro 's workshop on youtube (https://www.youtube.com/watch?v=CB3FKtmuPI8). The most important mismatch and the reason why I was getting NaN as a loss is inside: If this is included, then the script works! |
|
The other mismatches are (don't thinkt they are dramatic, but just so you know): batch_size=4, |
|
Hi @Atmoboran ! The ones I showed are actually here: https://github.com/IBM/terratorch/tree/organise_examples/examples/notebooks/Prithvi_EO We are in the process of reorg the examples and we did not manage to transition them yet! Sorry about it and thanks for reaching out! |
So I have found the problem while watching @paolo-fraccaro 's workshop on youtube (https://www.youtube.com/watch?v=CB3FKtmuPI8).
There are quite a few mismatches between the notebook he is showing and the one that can be downloaded from terratorch/examples/tutorial/prithvi_v2_eo_300_tl_unet_burnscars.ipynb
The most important mismatch and the reason why I was getting NaN as a loss is inside:
datamodule = terratorch.datamodules.GenericNonGeoSegmentationDataModule(
...
no_data_replace = 0,
no_label_replace = -1,
)
If this is included, then the script works!