Self-supervised learning (SSL) has become a stable for training visual foundation models from large amounts of unlabeled data. However, recent state-of-the-art methods, such as DINOv3, rely on increasingly complex training pipelines and design choices that are largely designed and evaluated with natural image datasets like ImageNet in mind. Many scientific imaging domains, including medical and pelagic imaging, violate these assumptions due to distribution imbalance, sparse informative content, and heterogeneous image statistics. As a result, current SSL approaches often struggle to learn representations that consistently outperform strong supervised baselines, despite the availability of abundant unlabeled data. This PhD aims to investigate how SSL methods can be made more robust and broadly applicable to non-natural image domains by identifying the limitations of current approaches and developing algorithms that better accommodate the characteristics of scientific datasets.