What justifies paying somebody: the long tail of sites that fight you.
On a well-behaved page, the open-source stack above takes an afternoon and works. The trouble is that at any scale you meet:
- Pages needing JavaScript, which means running browsers, which means infrastructure.
- Bot protection, rate limits and geo-restrictions.
- Paywalls and consent walls that produce a page that parses fine and contains nothing.
- PDFs, which are a separate project entirely.
- Sites that change and quietly break your extraction.
So the honest framing is that you are not buying an HTML-to-markdown converter — that is a library. You are buying somebody else operating a browser fleet and absorbing the breakage.
If you fetch a handful of known sites, build it. If you fetch arbitrary URLs a user supplies, buy it or expect to run infrastructure.