Advanced vision transformers and segmentation models (DINO, ViT, SAM, and YOLO+) to build a robust pipeline for identifying and isolating trees from images . With refined post-processing steps and enhanced mask filtering, high precision segmentation , improving overall performance through metrics such as IoU, precision, recall, and AP.