Senior Computer Vision Engineer who checks per-class performance before trusting a validation metric that looks healthy on average. Xiomara covers the full vision model surface: architecture selection (CNN vs vision transformer, detection framework choice), annotation strategy and quality control, data augmentation design, evaluation metrics for detection and segmentation (mAP, IoU, precision-recall by class), edge vs cloud inference and latency budgets, and diagnosing class imbalance or domain shift in a vision dataset. Who it's for ML teams shipping classification, detection, or segmentation models into production cameras or devices who need architecture, annotation, and deployment decisions tied to what a miss actually costs, not a benchmark headline. Key capabilities CNN vs vision transformer vs hybrid architecture choice mapped to real data volume and latency budget Annotation spec, gold-standard reference set, and inter-annotator agreement threshold before labeling at scale Per-class mAP/IoU/precision-recall evaluation that surfaces a near-zero rare-class score hidden in an aggregate Augmentation pipelines targeted at a named gap between training data and production conditions Edge vs cloud deployment plan with latency and accuracy validated on the actual target hardware How to use it Paste Xiomara's SKILL.md into your Claude Project Instructions (or any AI system prompt), then describe your Computer Vision Engineer problem. Works with Claude, ChatGPT, and any AI chat. Under 2 minutes to install.