Pipeline reports inherit whatever the contact records underneath them look like. I’d never actually measured that in a portal, so last week I did.
Pulled every contact carrying a LinkedIn URL through the search API:
POST /crm/v3/objects/contacts/search
{"filterGroups":[{"filters":[{"propertyName":"hs_linkedin_url","operator":"HAS_PROPERTY"}]}],"properties":["firstname","lastname","company","jobtitle","hs_linkedin_url","notes_last_contacted"],"limit":100}
Page on paging.next.after until you hit total.
1,836 records came back. Normalizing the URLs first (drop www, the country subdomain, query string, trailing slash) turned up 24 duplicates a plain string match had missed. 485 of the 1,812 uniques were non-US, which that subdomain tells you for free.
Then I opened 17 of the profiles and compared them to the records. Seven were a different person entirely. Two more listed employers they’d left years ago. A prospecting tool had built them by pairing a scraped name with whatever profile URL sat nearby.
All of those contacts associate to companies and deals. The forecast renders fine either way and nothing flags it.
Two things worth knowing if you check your own: notes_last_contacted beats lastmodifieddate for recency, because imports touch lastmodifieddate and make dead records look current. And large same-day batches sourced OFFLINE are usually imports worth sampling before anyone sells off them.
Happy to share the normalization logic if it’s useful.
We run free office hours on this stuff with Claude and HubSpot. Contact cleanup, deal diagnosis, the connector, API keys both directions.