Inference is what you do with the model. It's not an executable, so you don't run a model directly, instead a separate program runs and that program reads from the model. Also, hosting can be used to mean running inference on a model, but it could also be used to mean just storing the files and possibly making them available for download, so it's a bit more vague.
Inference sounds cooler and more mystical. Any old company can host something, "running inference" requires 10x rock star engineers and a CEO that acts like a badass who wears a leather jacket and has Thoughts about demographics.
That's interesting, the idea that hosting implies a front-end (whether UI or API).
So it's similar to why we don't call Heroku or Vercel "hosting" or "compute" because it's more of a service (though "cloud" and "platform" are pretty vague too).
I'd say "hosting" means making something available to you in the cloud. Model makers don't offer that.
They sell "inference," which is a compute task with real hardware costs for them. The task runs against their model, which is not hosted for you.
So it's similar to why we don't call Heroku or Vercel "hosting" or "compute" because it's more of a service (though "cloud" and "platform" are pretty vague too).