// Node.js MCP server · OAuth 2.1 + PKCE · Google Analytics / Search Console / Tag Manager / Cloudflare / DataForSEO

Google SEO & Site Ops MCP

A single authenticated MCP server that reads Google Analytics, Search Console and Tag Manager data, provisions new sites end to end (GA4 property, GSC verification, GTM container), manages DNS and email routing on Cloudflare, and runs SSRF-guarded competitive research captures with DataForSEO keyword data — 65 tools behind one owner-gated OAuth login.

Node.js 22 / Express 5 googleapis (GA4, GSC, GTM) OAuth 2.1 + PKCE, owner password Cloudflare API (DNS, Email Routing) playwright-core + Azure Blob DataForSEO

What It Is

This started as a reporting tool — ask an AI client for GA4 traffic, Search Console queries, or a conversion comparison across two periods — and grew into the thing that actually stands a new client site up end to end: create the GA4 property and stream, verify the domain in Search Console, submit the sitemap, create and publish the GTM container with its GA4 tag, point DNS at the host on Cloudflare, and set up inbound email forwarding to a verified mailbox, all through one MCP connection.

It's a single-owner server, not multi-tenant SaaS — one person (me) is the only principal who can ever authenticate to it, enforced by an owner password behind a standards-based OAuth 2.1 authorization-code-with-PKCE flow, or a static bearer token for simpler clients. Every one of its 65 tools, across six distinct domains — GA4/GSC/GTM reporting, site setup/provisioning, SEO research capture, DataForSEO keyword and competitor data, Cloudflare DNS/Email Routing, and diagnostics — sits behind that one gate.

The newest piece is a research module that captures a competitor's or prospect's public page with a headless browser (falling back to a plain HTTP fetch when no browser binary is available), extracts structured evidence — headings, forms, links, meta tags — and explicitly tags anything unverifiable (testimonials, structured-data claims) rather than presenting scraped marketing copy as fact.

65 Tools, Six Domains

DomainTool countExamples
GA4 / GSC / GTM reporting15ga_run_report, ga_compare_period, gsc_search_analytics, gsc_page_queries, ga_top_pages
Google site setup/provisioning16google_setup_accounts, ga_setup_site, gsc_prepare_verification, gsc_verify_site, gtm_setup_container, gtm_publish_version
Niche research (capture/scoring)8research_start, research_capture_page, research_score_page, research_build_brief
DataForSEO keyword/competitor/backlink13research_keyword_volume, research_keyword_difficulty, research_competitor_keywords, research_backlinks
Cloudflare domains/DNS/Email Routing13cloudflare_list_zones, cloudflare_create_dns_record, cloudflare_enable_email_routing, cloudflare_create_email_forward
Operability—Health check, config/diagnostic status

Tech Stack

Server runtime
Node.js 22 (ESM)type: module
Express 5HTTP + OAuth routes
@modelcontextprotocol/sdkStreamable HTTP + legacy SSE transport
zod + zod-to-json-schemaTool input validation
Google platform
googleapisAnalytics Admin + Data, Search Console, Tag Manager, Site Verification
OAuth2 refresh token orservice account JSON (scope-gated by write flag)
Research / capture
playwright-coreHeadless capture, no bundled Chromium
cheerioServer-side DOM parsing of captured HTML
undiciLow-level fetch with DNS pinning for SSRF defence
ipaddr.jsPublic vs. private/reserved address classification
Infra / integrations
@azure/storage-blobHosted evidence/screenshots
Cloudflare API v4Zones, DNS records, Email Routing
DataForSEO APIKeyword volume/difficulty, competitor, backlinks
Azure App ServiceHost for the remote MCP endpoint

Single-Owner OAuth 2.1 + PKCE

Every MCP transport (/mcp, /sse, /messages) sits behind one middleware. There's no multi-user model here by design — this authenticates one person to one server, standards-compliant enough that any MCP client can complete the flow itself.

client registers via /oauth/register→ owner enters password at /oauth/authorize→ PKCE-verified code exchange at /oauth/token→ signed access + refresh token
A

Stateless, HMAC-signed tokens when MCP_TOKEN_SECRET is set

Client registrations, access tokens and refresh tokens are all signed with HMAC-SHA256 and verified on every request rather than looked up from memory — so an App Service restart or redeploy never logs a connected client out. Without the secret configured, the server falls back to the original in-memory behaviour (reconnect after a restart). Rotating the secret instantly revokes every outstanding token, which is the actual "kill switch" if a token ever leaked.

B

PKCE is mandatory, not optional

The authorization endpoint requires an S256 code challenge matching the exact 43-character base64url format before accepting a request at all, and the token endpoint independently re-verifies the code verifier's SHA-256 digest against the stored challenge before ever issuing a token — standard OAuth 2.1 practice, enforced strictly rather than treated as a client's responsibility to get right.

C

CSRF-bound login, scoped Content-Security-Policy, timing-safe comparisons

The password-entry page sets an HttpOnly, SameSite=Lax session cookie tied to that specific authorization request; the POST back must present the matching cookie before the password is even compared. Every secret comparison (password, CSRF token, HMAC signature) runs through a constant-time timingSafeEqual wrapper over SHA-256 digests, specifically to avoid a timing side-channel on any of these checks. The authorize page's own CSP is scoped to exactly the registered redirect origins — form-action 'self' <allowed origins> — so the login form itself can't be repurposed to post credentials somewhere unintended.

D

A static bearer token path exists for clients that can't do OAuth

MCP_ACCESS_TOKEN (32+ characters, compared with the same timing-safe check) is accepted as an alternative to a session-issued token — useful for a simple script or a client that doesn't implement the full authorization-code flow — without weakening the OAuth path for clients that do.

Architecture & Design Decisions

01

SSRF-safe page capture — DNS-pinned fetch, not a trusting one

Before capturing any URL, resolvePublic() resolves its hostname and rejects the request outright if any resolved address is private, loopback, link-local or otherwise reserved (via ipaddr.js range classification) — a prospect's "look at this page" request can't be used to make the server fetch its own internal network. The actual fetch then pins the socket to that exact resolved address with a custom undici dispatcher rather than re-resolving DNS at connect time, closing the classic DNS-rebinding gap where a hostname resolves safely at check-time and unsafely at connect-time. Redirects are capped at five hops and each hop is independently re-validated through the same public-address check.

02

Graceful degradation from full browser capture to plain HTML fetch

Capture first looks for a real Chromium binary across several known paths (bundled Playwright browser, system Chromium, Edge/Chrome on Windows); if none exists, it falls back to the plain DNS-pinned HTTP fetch and extracts whatever static HTML evidence it can, explicitly flagging the degraded result's limitations ("JavaScript content and visual/mobile usability are unassessed") rather than silently returning a thinner result as if it were equivalent to a full capture.

03

Evidence extraction that tags its own unreliable categories

extractEvidence() pulls titles, meta description, headings, forms (with field types and resolved labels) and links into a structured evidence list — but anything from a page's own structured data (JSON-LD) is tagged structured_data_claim_unverified, and every result carries an explicit limitations array stating that reviews, qualifications and other on-page claims aren't independently verified, and that no conversion-rate, backlink-authority or Core Web Vitals measurement happened. The tool is built to make it hard to accidentally present a scraped claim as a verified fact downstream.

04

Content-hash optimistic concurrency on Cloudflare DNS writes

Every DNS record read is returned with a recordVersion — a SHA-256 hash of the record's own canonicalized (key-sorted) fields — and a write tool can require that exact version to match before applying an edit. If someone (or something else) changed that record between the read and the write, the hash won't match and the edit is rejected rather than silently overwriting a change neither side knew about. Record validation is also strict before any API call: an A record's content must parse as a real IPv4 address, a CNAME's target can't itself be an IP or contain a wildcard, and a proxied record is forced to automatic TTL — Cloudflare's own API would reject some of these, but failing fast locally gives a clearer error than relayed API error codes would.

05

Two Google auth modes, each scoped to the minimum it needs

createGoogleAuth() supports either a user OAuth refresh token (all three of client ID, secret and refresh token must be present together — no silently-partial config) or a service-account JSON credential. Read-only reporting calls request narrow analytics.readonly / webmasters.readonly scopes; only the explicit "setup" path that provisions new GA4 properties, GTM containers and GSC verification requests the broader write scopes (analytics.edit, tagmanager.edit.containers, tagmanager.publish, etc) — a reporting-only client session never holds write-capable Google credentials at all.

06

Provisioning never silently retries a resource-creation call

The setup Google API clients are explicitly configured with retry: false — a deliberate choice, because automatically retrying an ambiguous failure on a POST that creates a GA4 property or publishes a GTM container risks creating a duplicate resource rather than just duplicating a safe read. The provisioner also uses an only() helper that refuses to proceed when a lookup matches more than one existing resource, forcing an explicit ID instead of guessing which one the caller meant.

07

Sitewide default: every new site gets an owner-verified email forward

The provisioning workflow sets up info@<domain> forwarding to the verified owner mailbox for every new site by default (overridable), rather than leaving a freshly provisioned domain with no working contact address — a small operational detail baked into the automation specifically because it was the kind of thing that was easy to forget doing manually for every new client site.

The Timing-Safe Comparison Pattern

// access-control.js — every secret comparison in the server goes through this const digest = value => createHash('sha256').update(String(value)).digest(); const equal = (a, b) => typeof a === 'string' && typeof b === 'string' && timingSafeEqual(digest(a), digest(b));

Q&A

Q
Why build a custom OAuth flow instead of just using a static API key for everything?
Some MCP clients (Claude's connector flow among them) expect standards-based OAuth 2.1 discovery — the /.well-known/oauth-authorization-server and protected-resource metadata endpoints — to connect at all, so a static key alone wasn't enough for every client I wanted to support. I kept the static bearer token path too, specifically for simpler scripts that don't need the full dance, so neither approach is the only option.
Q
How do stateless signed tokens actually work, and why did you need them?
Early on, every App Service restart or redeploy logged my connected AI client out, because tokens lived only in server memory. With MCP_TOKEN_SECRET set, a token is just a base64url JSON payload plus an HMAC-SHA256 signature over it — verified by recomputing the HMAC and checking expiry, nothing stored server-side at all. Restarts don't affect it because there's no state to lose; rotating the secret instantly invalidates everything, which is the recovery path if a token ever leaked.
Q
Walk me through the SSRF protection in the research capture tool.
Before fetching any URL, I resolve its hostname and reject it if any resolved address is private, loopback or otherwise non-public — using ipaddr.js's range classification, not a hand-rolled IP regex. Then the actual HTTP request pins the connection to that exact validated address with a custom undici dispatcher, rather than letting the HTTP client re-resolve DNS when it actually connects — that gap, a hostname resolving safely at check time and differently at connect time, is a known SSRF bypass (DNS rebinding), and pinning the socket closes it.
Q
How does the Cloudflare DNS tool prevent a write from clobbering someone else's change?
Every record read returns a content hash of its own canonicalized fields as a recordVersion. A write tool can require that exact hash to still match before it applies the edit, so if the record changed between my read and my write, the hash mismatches and the write is rejected rather than silently overwriting whatever changed it. It's the same optimistic-concurrency idea as an ETag, implemented as a plain SHA-256 hash because Cloudflare's API doesn't give me one natively.
Q
Why keep this single-owner instead of building it multi-tenant like EventQuoter?
This tool operates directly on my own Google Analytics, Search Console, GTM, Cloudflare and DataForSEO accounts — provisioning real GA4 properties and publishing real GTM containers. There's no case here for a second principal; the entire security model is built around exactly one owner, which is why the password-gated OAuth flow and the single MCP_ACCESS_TOKEN fallback are both designed around "exactly one person can ever get in," not "many tenants, properly isolated."
Q
What would you improve next?
The in-memory fallback path (when MCP_TOKEN_SECRET isn't set) still loses connected clients on every restart — fine for local development, less fine if someone forgets to set the secret in production. I'd also want to extend the DataForSEO research tools with scheduled tracking over time rather than point-in-time snapshots, so keyword-position and backlink changes become visible as a trend rather than something I have to re-run manually to notice.