Hi there,
Our server is getting requests to private URLs with a User Agent of: “HubSpot Links Crawler 2.0 http://www.hubspot.com/”
It cannot access these URLs and I understand that it will respect robots.txt if provided.
What I’d like to understand is how HubSpot is getting these URLs to scrape in the first place. They should only be available from private pages, HubSpot shouldn’t know these URLs exist.
Is there anything available to help understand how HubSpot is getting these URLs to scrape?
Thanks,
Darin