We’ve not had issues getting T4 GPU instances from AWS. We faced difficulties provisioning A100 and AWS is annoying that the 2 tiers are either T4 or directly A100’s I think.
We use AWS GPU machines for our CI but for serious ML training workloads we use GCP L4 GPU instances. Even in GCP we couldn’t provision or A or H100 (our quota itself is just 1 GPU of these instances) but we’ve never had issues provisioning L4 GPU’s and I think that’s enough for smaller not LLM Scale Models. For LLM scale startups, it’s tough provisioning GPU’s even if you have money.
We use AWS GPU machines for our CI but for serious ML training workloads we use GCP L4 GPU instances. Even in GCP we couldn’t provision or A or H100 (our quota itself is just 1 GPU of these instances) but we’ve never had issues provisioning L4 GPU’s and I think that’s enough for smaller not LLM Scale Models. For LLM scale startups, it’s tough provisioning GPU’s even if you have money.
(We’re based in Bay Area btw)