See the isolation
The query has no tenant filter. Row-Level Security still returns only the requesting tenant's rows. Simplified illustration with sample rows; the policy shape is ComplyDesk's.
-- request from acme.complydesk.online SELECT set_config('app.tenant_id', 'acme', true); SELECT * FROM "Evidence"; -- no WHERE tenantId
What it is
Small companies chasing SOC 2 mostly run compliance on a spreadsheet: a list of controls, a folder of screenshots labelled “evidence,” and someone’s memory of what’s expired. ComplyDesk replaces that spreadsheet with a multi-tenant register. Each company gets a workspace seeded with a standard control set, uploads evidence against each control, and gets a dashboard of what’s covered, what’s missing, and what’s about to expire.
An AI layer on top reads uploaded evidence well enough to auto-classify it against a control, and reads an uploaded policy document well enough to flag which controls it doesn’t actually cover.
The demo tenant is real, pre-generated data: an 18-control set, evidence classified by the actual Gemini pipeline, a policy document chunked and embedded, and a gap-analysis report from a real Groq run with real citations. One file is marked FAILED because it genuinely tripped Gemini’s free-tier rate limit mid-seed, and was left as-is rather than faked. None of it was hand-inserted into the database. Uploads and re-runs are switched off on that one tenant by a server-side guard, so one visitor can’t spoil the next visitor’s tour.
Stack
- Frontend
- Next.js 16 (App Router), React 19, on Vercel
- API
- NestJS 11 on Render. CommonJS and Jest, because Nest 12 ships ESM-only, which classic Jest can’t consume
- Database
- PostgreSQL on Neon, via Prisma 7 with the
pgdriver adapter - Vector search
- pgvector,
ivfflatindex, 384-dim embeddings - Object storage
- Neon Object Storage (S3-compatible), presigned URLs only
- AI
- Gemini (vision and embeddings) and Groq (text reasoning) behind one interface
- Monorepo
- npm workspaces:
apps/web,apps/api,packages/shared
Request path
Every request to a tenant subdomain such as acme.complydesk.online goes through the same sequence before any handler runs.
Resolve the tenant
TenantMiddlewarereads the Host subdomain, falling back to anX-Tenant-Slugheader, and stores the tenant id in request-scoped context.Open a scoped transaction
TenantTransactionMiddlewareopens one Prisma transaction for the rest of the request and runsset_config('app.tenant_id', id, true), which hasSET LOCALsemantics.Carry context without prop-drilling
nestjs-cls(AsyncLocalStorage) carries the tenant, user, role and transaction through every guard, interceptor and handler.Authenticate, authorise, audit
JwtAuthGuard, thenRolesGuard, thenAuditInterceptor, which records every mutation in the same transaction as the write.Hit Postgres as a constrained role
Queries run as
app_runtime, aNOBYPASSRLSrole, against tables withFORCE ROW LEVEL SECURITYand policies of the formtenantId = current_setting('app.tenant_id').
A service method that forgets to add WHERE tenantId = … still doesn’t leak. RLS filters it at the database, regardless of what the application code remembered to write.
Tenant isolation
The whole design rests on one claim: an application bug in a service method cannot leak another tenant’s row. Getting from “we usually filter by tenant” to something I can stand behind took a few Postgres and Neon specifics most CRUD apps never touch.
FORCE ROW LEVEL SECURITY, and the role it binds
ENABLE ROW LEVEL SECURITY alone doesn’t restrict a table’s owner, which quietly defeats the point if the app connects as the same role that ran the migrations. The app connects as a separate least-privilege app_runtime role instead: LOGIN, NOBYPASSRLS, DML on exactly the tables it needs, no DDL. Migrations run as the owner.
Two Neon-specific footguns surfaced along the way, both documented inline in the migration. Roles created through Neon’s console or CLI get BYPASSRLS by default, so app_runtime is created in plain SQL instead. And CREATE ROLE is cluster-wide, so Prisma’s shadow database creates the role for real before the migration reaches the target. Every statement in the RLS migrations is idempotent for that reason.
SET LOCAL, not SET
Tenant scope is set with set_config(…, true), which is transaction-local. A plain session-wide SET would persist on the pooled connection after commit, and the next request handed that connection would silently inherit the previous tenant’s context. The value is passed as a bound parameter, since this is the one place tenant id touches raw SQL.
The ::uuid cast is load-bearing
On a fresh connection, an unset app.tenant_id makes current_setting() throw, which is what should happen. On a pooled connection where an earlier transaction already set it, Postgres returns an empty string instead. A text comparison against '' would read as “this tenant has zero rows”: a wiring bug that looks exactly like an empty account. Casting to ::uuid turns it back into a loud error. Both behaviours are asserted in the test suite, not just reasoned about in a comment.
Tenant and User need different policies
Neither has a tenantId column. Tenant leaves SELECT open deliberately, since the middleware has to look a tenant up by slug before any tenant context exists, while writes stay scoped with both USING and WITH CHECK. User is scoped through an EXISTS subquery on Membership. Testing on a Neon branch showed a qualified cross-tenant UPDATE was already blocked by the SELECT policy alone. The gap the membership policy actually closes is the unqualified bulk write: it affected 2 rows across 2 tenants under a naive policy, and exactly 1 under the scoped one.
Tested adversarially
Roughly 30 cases in rls-isolation.e2e-spec.ts, run against a real Neon branch with no mocked database: cross-tenant SELECT, UPDATE and DELETE, WITH CHECK forgery on INSERT, a service method with no tenant filter in its own code, the pooled-empty-string throw, and per-table coverage of every RLS-protected table, including the five AI-layer tables added later.
Other decisions
Two AI providers, split by what each is good at
Evidence classification (which needs vision, for photos and scans) and embeddings stay on Gemini. Gap analysis, which is pure text reasoning over policy chunks, moved to Groq. Both sit behind one AiProvider interface; callers don’t know two vendors are involved. The move happened because a single gap-analysis run cost about 18 of Gemini’s roughly 20 free requests a day. The first Groq model I picked turned out to be retired, found through 404s on a real run rather than by reading changelogs. After switching to a model confirmed in the account’s live list, a real run finished in about 51 seconds with zero errors and zero 429s, and the rate limit was re-derived from that data.
Being honest about tenant resolution
Browsers never let JavaScript override the Host header, so the web app sends X-Tenant-Slug as a fallback. That isn’t a security hole on its own, because JwtAuthGuard still requires a real membership for the resolved tenant, but it does mean tenant resolution trusts a client-supplied value. The production answer is a reverse proxy that forwards the real subdomain. I left it out on purpose to keep scope on the multi-tenancy and RLS work.
Audit logging inside the request’s transaction
AuditInterceptor records every POST, PUT, PATCH and DELETE with actor, action, target and a before/after diff, through the same transaction the handler writes with. The audit row commits or rolls back with the mutation, and it inherits RLS, so it cannot write into another tenant’s log. It’s best-effort by design: an audit failure is logged and never allowed to fail the original request.
Lessons learned
- Login failed right after signup, but only when email case differed. The cause was that
User.emailhad no case normalisation and Postgres text equality is case-sensitive, not bcrypt or the new RLS policies. Fixed at the DTO boundary and independently with aCHECK (email = lower(email))constraint. Chasing it surfaced a second bug: no e2e test had been running through the app’s realValidationPipe, so DTO validation had been silently inert in every test. nest buildproduced an emptydist/with exit code 0. The incremental build cache lived outsidedist/, so afterdist/was wiped, tsc trusted the stale cache and compiled nothing. Fixed by moving the cache insidedist/./auth/mereturned 403 with a valid token. Host-based resolution can never see a tenant on a browser request, which is whyX-Tenant-Slugexists. The same investigation found CORS rejecting every tenant subdomain, and apatternattribute that Chrome’s regex engine rejected, breaking the signup slug field in one browser only.- A Render build failed with about 37 type errors that never appeared locally, because the build never ran
prisma generateand Render starts from a clean container. I reproduced it locally by deleting the generated client, then fixed the build script. - Deploying the web app to Vercel from inside
apps/webuploaded only that directory, with nopackages/shared. Fixed with a root-levelvercel.jsonand explicit build settings.
What I’d do next
- Replace the client-supplied
X-Tenant-Slugfallback with a reverse proxy that forwards a trusted header. - Move rate limiting to Redis.
TenantRateLimitGuardis an in-memory fixed window scoped to one process, which is a known and accepted limit for now. - Take a deliberate second look at the
TenantandUserRLS policies before either table grows new tenant-facing surface, since their shape differs from every other table’s. - Replace the hand-rolled root
vercel.jsonwith a proper Root Directory setting once the CLI supports it non-interactively.
- Next.js 16
- React 19
- NestJS 11
- PostgreSQL
- Prisma 7
- pgvector
- Gemini
- Groq
