Skip to content
Fabian Finalé Franqui
Selected work

Case study 01

Building an enterprise integration across frontend, identity, and platform boundaries

The web architecture for an onboarding experience launched from an external enterprise platform: a server-driven workflow, identity integration, asynchronous account setup, and observability by design — delivered as the primary frontend engineer, across several repositories.

My role
Primary frontend engineer: web architecture, implementation, and instrumentation
Scope
Web funnel, sign-up surface, backend-for-frontend endpoints, targeted identity changes, telemetry and dashboards
Focus
  • Frontend architecture
  • Enterprise integration
  • Observability
  • AI-assisted engineering
Validated returnSigned inEnterprise hostExternalWeb appEntry · funnelIdentity platformAuth · sessionSign-up / loginAccount creationTelemetryBackend-for-frontendFunnel endpointsBackend servicesRules · verificationObservabilityEvents · dashboards
Fig. 1The integration and its boundaries. The host launches the web app — designed for an embedded frame, launched as a full-page flow; sign-in runs through the identity platform and a dedicated sign-up surface before returning to the funnel, which reaches backend services through a backend-for-frontend; every screen reports telemetry. Highlighted: the parts I owned or led.
~4 months
Of build, after an earlier design phase
Multi-repository
Web app, sign-up surface, backend-for-frontend, and identity layers
19 instrumented screens
Views, steps, blocks, and interactions, with dashboards built on them
Fail closed
Eligibility, unknown workflow states, and the return path default to the safe side

Context

An external enterprise platform wanted to offer a financial product to its users from inside its own application. A user would arrive from the host, see a value proposition, create an account, verify their identity, accept terms, wait while the account was set up by an asynchronous backend operation, and return to where they had started. It was designed and built to run inside an embedded frame in the host’s application; late in the project, the host’s security constraints changed, and it launched as a full-page flow instead.

On our side, that meant an integration across several systems and teams: the web application, a separate sign-up and login surface, the identity platform, a backend-for-frontend, the backend services for the workflow and for identity verification, and the analytics pipeline. I came in as a frontend engineer and became the primary one for the web experience through most of the build.

Problem

The hard parts were at the seams. An experience embedded in another company’s application runs as a third party in the browser, which changes what cookies, storage, and sessions can be relied on. The sequence of screens depended on state only the backend knew — eligibility, verification outcomes, whether the user had already finished — so any order hard-coded in the frontend was a guess. Account setup depended on an operation that could be slow or fail. And every one of those paths had to be visible after launch, not reconstructed from support tickets.

None of it could be solved in one codebase. The frontend could only be correct if the backend-for-frontend, the identity platform, and the backend’s workflow rules agreed with it, and most of those belonged to other teams.

Constraints

  • The experience was designed with the host to run inside an embedded frame, with the host deciding what the user could do and how they could leave.
  • The identity and backend platforms were owned by other teams; the integration depended on their contracts and their release schedules.
  • Configuration the flow depended on — redirect destinations, allowed origins, contract versions — lived in several repositories at once.
  • Eligibility and compliance outcomes were enforced server-side; the frontend had to present every one of them, and fail closed on any it didn’t recognize.
  • Behavior and analytics had to stay comparable with the native mobile apps, built in parallel by another team.
  • Local end-to-end testing was unreliable for much of the project, and the shared mock server was decommissioned midway through.

My role

I was the primary frontend engineer through most of the build, responsible for the majority of the frontend changes merged in scope. I wrote the web engineering design and carried it through review, and I owned the funnel’s navigation and eligibility handling, the embedding and session hardening on the web side, the test infrastructure, the analytics instrumentation, and the backend-for-frontend endpoints the funnel needed. Where the integration required it, I made targeted changes in the identity layer, the backend workflow, and the verification service, and I took over most of the sign-up surface’s changes once its foundation was in place. Other frontend engineers built the first screens and that foundation; the identity, backend, and mobile teams owned their platforms; and shipping the integration was a shared outcome.

Decisions

  1. 01Let the server drive the workflow

    The first version hard-coded the order of screens. It was replaced with navigation driven by the backend’s workflow state: each step reports whether it is complete and whether it should be shown, and a single map from step to route decides where the user goes next. The frontend stopped guessing and started reflecting.

    TradeoffCorrectness now depends on the backend’s step semantics, which had to be made explicit: only the workflow’s top-level status says a user is finished, and only at entry. Written down, it still got re-derived from scratch once.

  2. 02Classify returning users at entry

    Whether someone had already finished had to be decided once, when they arrived, from the one signal that was reliable there — and never again mid-flow, where the same state was ambiguous. An earlier attempt to make that call mid-flow caused a regression that had to be reverted, and settled the question.

    TradeoffA dedicated entry route is one more path to test and instrument. It is also the only place the answer can be trusted.

  3. 03Fail closed on eligibility

    Eligibility reasons from the backend map to specific block screens through an explicit allowlist. Anything unrecognized blocks rather than proceeds, a re-check runs after identity verification, and every block shares one guard.

    TradeoffFailing closed can over-block when the backend adds a status the frontend hasn’t learned yet. For a compliance-sensitive flow that is the right side to err on, and the block screens are instrumented, so it shows up in the data.

  4. 04Treat embedding as a cross-layer problem

    Cookies that were dropped, storage access that threw, request protections that failed without a log line: each looked like a frontend bug and had its root in another layer. Rather than work around them in browser code, the fixes went where the cause was — guards in the frontend, cookie handling in the backend-for-frontend, session attributes and handoff in the identity layer — with security review on the sensitive parts.

    TradeoffIt meant changes in services other teams owned, and one of mine was reverted. Fixing the root layer was still cheaper than carrying workarounds in the frontend indefinitely.

  5. 05Move to a full-page flow when the host’s constraints changed

    The embedded design was agreed with the host from the start and hardened for weeks. Late in the project, the host’s security constraints changed — its webviews could not allow third-party cookies, and Safari’s engine restricted them further — so the experience launched as a full-page flow, on its original date.

    TradeoffA change of that kind, that late, is the one that moves launches. The lesson is not that the question went unasked, because it was asked at the start. It is that a host’s answer can change, and the integration has to be able to change with it.

  6. 06Make the exit a single validated path

    The way back to the host changed three times as the host’s constraints became clear. Whatever the mechanism, the destination passes through one validated, allow-listed sink and survives the sign-in resets along the way.

    TradeoffAn allow-list is one more shared string across repositories, and it has to stay in step with the host’s environments.

  7. 07Defer telemetry until identity has resolved

    Screen views that fired before the user’s identity was known were duplicated and misattributed once it arrived. Page tracking now waits for a valid identity, and screens that can re-render latch their first view.

    TradeoffA screen abandoned before identity resolves never reports a view. That loss is bounded and understood; misattributed data is neither.

  8. 08Ship in small, stacked changes

    The first implementation arrived as one large change and was closed unmerged; it came back as a stack of six, each reviewable on its own. The same pattern carried the exit re-architecture and the analytics work.

    TradeoffMany small changes inflate every count and make the history noisier. Reviews move faster, and a revert takes one change with it instead of six.

Server-driven workflow

The original routing encoded a sequence: value proposition, sign-up, verification, terms, setup, done. It also encoded an assumption — that the signal it used to detect an existing user meant the same thing for everyone. It didn’t, and users who had already finished were routed back into identity verification.

Moving the sequence to the server changed the shape of the problem rather than its size. The backend’s workflow became the source of truth for which step comes next; the frontend’s job became presenting that step correctly, including the states where there is no next step: ineligible, under review, already enrolled, or a status nobody had planned for. The invariants that make this safe — finished means finished only at entry; unknown means blocked — are the kind of thing that has to be written down, because the fixtures that stood in for the backend didn’t enforce them, and they hid the regression until real data found it.

What the frontend had to handle

  • Returning users, identified once at entry and sent to a dedicated screen instead of back into the funnel.
  • Verification outcomes: a pass, a soft fail with retry guidance, a hard fail with an exit, and a manual review with its own screen.
  • Eligibility blocks for users the product couldn’t serve, each with its own screen and its own event.
  • An explicit fail-closed default for any status the frontend didn’t recognize.

Embedding and identity

Running inside another company’s application makes the experience a third party in the user’s browser. Cookies the application relied on were dropped or unreadable; storage access could throw; request protections failed with a bare status code and nothing in the logs; and the host decided which navigation was allowed at all. Each symptom surfaced in the frontend, and almost none of them could be fixed there.

The work crossed three layers. In the browser: guards around storage, a clean-slate sign-in to defeat stale sessions, and a forced refresh at the identity provider when the signed-in user didn’t match the one the host had sent. In the backend-for-frontend: cookie handling for embedded origins. In the identity layer: session attributes for the embedded context, redirect registrations per environment, and skipping sign-up steps that didn’t apply to the host’s flow. I made the narrow changes in those services myself, with security review, and worked with the identity team on the parts they owned — including the session reset earlier in the flow that became the real fix after my own attempt was reverted.

Identity itself stayed with the identity team: the client and flow design, registration, session resets, and the versioned contract between the surfaces. My side was the integration — selecting the right client for the host, handling the entry token the host mints, including a reuse problem I traced through browser request captures, and keeping the return destination alive through every reset along the way. Several identity-side failures I diagnosed from the frontend by reading the identity extensions’ source, and fixed the ones that were mine to fix.

The exit tells part of the story. It began as a close button, was hidden, then removed once the host declared itself the owner of the exit. The return went from a redirect, to a message to the parent frame — six weeks on that mechanism — and back to a redirect. The bigger change came late in the project: the host’s security constraints changed, its webviews could not allow third-party cookies at all, and Safari’s engine restricted them further, so the experience launched as a full-page flow. By then the move was small — the return buttons that had messaged the parent frame to close its modal became a validated, allow-listed full-page redirect — and the launch held its date. The frame had been discussed with the host from the first conversation; what changed was the answer. An embedded frontend problem is rarely a frontend problem, and sometimes it stops being an embedding problem at all.

Asynchronous setup

Creating the account triggers a backend operation that depends on another system and can take a while, or fail. The screen that waits for it looks simple and carried most of the reliability bugs: a polling effect that stalled because its dependency never changed by reference; a shared cache key that let a stale workflow read as success; returning users stuck on “setting up” until their identification moved to entry.

The wait state ended up with a state-based attempt counter, cache invalidation around the operation, a longer polling window with escalating intervals, a “taking longer than usual” variant, and — when the operation failed — a retry with a bounded budget before a failure screen. The loader is held until the operation has genuinely settled, because an optimistic success here produces an account that looks ready and isn’t.

Observability by design

Telemetry was part of the product, not a follow-up. Every funnel screen reports a view with a shared envelope — product, source, a stable identifier for the host, and the workflow step — and each screen adds its own type, flow, and context. Steps on the advance path report an enrollment-step event under the backend’s real step names, so conversion by step reads the same in the dashboard as in the workflow. Block and exit screens report why a user stopped. Interactions report what was tapped and where it led. Verification-failure screens mirror the properties the mobile apps already sent, so the two platforms can be compared on one chart. A compliance-sensitive gate was instrumented to report its outcome and attempt count to the logging platform as well, behind a flag of its own.

The first version got the ordering wrong. Screens that gained shared properties re-fired their view when the host identifier resolved, so the dashboards double-counted and misattributed views. The fix was to defer page tracking until identity had resolved and to latch the first view on screens that re-render — and to write the convention down, after I had first documented a naming rule the wrong way round and had to correct it. Two analytics golden files check in CI that the critical legs still emit what they should; I flagged alongside that work that they are blind to new properties, and that one pre-authentication drop-off is invisible to the step event.

On top of the events, I extended the team’s funnel dashboard with per-screen census, exits, interactions, and property coverage, with matching development and production views.

Questions the instrumentation answers

  • Where users from the host drop off, screen by screen, and which block or failure they hit.
  • How new users split from returning ones, and what the consent gate costs.
  • The mix of verification outcomes, and how often setup fails and is retried.
  • Conversion by workflow step, comparable with the mobile apps.

Testing and accessibility

Every page, block modal, hook, and route decision has unit tests (Jest with React Testing Library), and the critical screens have browser-level component tests with visual snapshots (Playwright), running against request-level mocks of the workflow — the same fixtures that drive local development. Page objects keep those tests readable, and a recorded video and trace of a run became the evidence attached to reviews. What I did not have was a reliable end-to-end run against a real backend: local end-to-end was unreliable for months, and the shared mock server was decommissioned midway through. When a ticket asked me to rebuild it, I pushed back and replaced it with request-level fixtures instead. The gaps are known: the snapshot tolerance can hide a copy change, and blocked states are covered only at the unit level.

Accessibility came back from the host’s screen-reader review as a list, and became a focused two-week remediation: a route-change announcer, headings and live regions per screen, keyboard-operable disclosures, label and id fixes, and several fixes in the shared design system — toast announcements, help text, focus order — that every other flow built on it inherits.

AI-assisted engineering

I used an AI coding assistant throughout, the way the deep dive on spec-driven development describes: as leverage, with the judgment kept on my side. Across several repositories, it was the fastest way to learn an unfamiliar part of a codebase, trace a failure, draft a test, or decompose a change into a reviewable stack.

Where it earned its place

  • Discovery: navigating unfamiliar services, the identity extensions included, to find where a behavior actually lived.
  • Investigation: root-causing the entry-token reuse problem from browser request captures, and building reusable recipes for log and analytics queries.
  • Repetition: test scaffolding, parity audits between web and mobile, snapshot generation, and the mechanics of stacked changes.
  • Analytics: constructing dashboard charts from the event definitions.

Where it was wrong, and what caught it

  • Its first proposed fix for the token problem would have dropped the host’s context from the session. It never shipped: the proposal was checked against the real contract first.
  • Stale checkouts of sibling repositories led it to confident parity conclusions that were simply out of date. Syncing before trusting became a rule.
  • It recorded an analytics naming convention the wrong way round, and the note had to be corrected later.
  • Edits landed in the wrong repository or working tree more than once, quietly voiding a test run. Checking where a command runs is now part of the loop.
  • It re-derived, from scratch, an invariant that was already documented — the same mistake a person makes when the document isn’t read.

None of this changed who was accountable. Scope, trade-offs, the security-sensitive cookie and session changes, which fixes to land, and the decision to split work into stacks were mine; every generated change went through the same review and tests as any other. The assistant materially accelerated discovery-heavy, repetitive, and investigative work. I don’t have a clean counterfactual, so I won’t put a multiplier on it.

What changed along the way

A four-month integration across teams does not go to plan, and the reworks are part of the record.

  • One change became six

    The first implementation was a single change far too large to review. It was closed unmerged and re-cut into a stack of six, which set the pattern for the rest of the project.

  • A routing fix caused a regression

    Classifying returning users mid-flow looked right in the fixtures and broke against real data. It was reverted, the classification moved to entry, and the invariant was written down.

  • A session fix was reverted

    My fix for stale sessions in the identity layer caused server errors and was reverted within two days. The identity team’s approach, earlier in the flow, became the fix.

  • The exit was designed three times

    Close button, parent-frame message, top-level redirect — each driven by a constraint that only became clear once the previous one was built. The dead exit modal was deleted later.

  • The frame gave way to a full-page flow

    Agreed with the host from the start and hardened for weeks, the embedded frame was dropped late in the project when the host’s security constraints changed. The move was small by then — a parent-frame message became a validated full-page redirect — and the launch held its date.

  • Duplicated configuration bit back

    Redirect destinations, allowed origins, and contract versions were duplicated across at least three repositories. A change that split one list by environment broke sign-in in development; I diagnosed and fixed it, but the duplication itself was never systematically solved.

  • The gating flag became a dependency

    The remote flag that gated the whole experience failed closed when its configuration couldn’t load, turning away users it should have let in. It was removed from the code, with the experience on for everyone.

  • The design document fell behind

    The web engineering design wasn’t updated after the first weeks, although the routes, the exit, and the shared configuration all changed after it was written. It is accurate about the tenets and wrong about several details.

Phases

  1. Design

    The web engineering design — tenets, scope, boundaries, open questions — written and reviewed while the backend and mobile designs ran in parallel.

  2. Foundation

    The first screens and the funnel shell, delivered as a stack, with a sandbox environment for the host’s testing at the first milestone.

  3. Identity and verification

    The host-specific sign-in client, entry-token handling, and the verification screens, with the backend-for-frontend and verification-service changes they needed.

  4. Server-driven funnel

    Backend-for-frontend endpoints for the workflow, navigation driven by its steps, and returning users kept out of the funnel.

  5. Hardening and rework

    Cookies, sessions, and storage across the three layers; copy and layout fixes from the host’s testing; the exit redesigned.

  6. Accessibility and eligibility

    The two-week accessibility pass, then the block screens, verification outcomes, and returning-user screens.

  7. Observability

    The event foundation, step events, deduplication guard, and the dashboards.

  8. Full-page flow

    Late in the project the host’s constraints change; the parent-frame message becomes a validated full-page redirect, and the launch date holds.

  9. Production enablement

    Production identity registration by the identity team, a redirect regression found and fixed, and the gating flag turned on fully and then removed.

Outcome

The experience was enabled in production on its planned date: the flag that had gated it was turned on for everyone and then removed from the code, with the production identity registration completed by the identity team. The web scope I owned was delivered: the server-driven funnel, the verification screens, the eligibility blocks, the setup wait state, the validated return path, the sign-up surface changes, the accessibility remediation, and the instrumentation.

What the team has now is a flow that can be reasoned about: the backend decides, the frontend reflects, every screen and block is measured, and the dashboard can answer where users stop and why. Adoption and completion numbers belong to the business, so they stay out of this write-up.

Lessons

  • Server-driven navigation moves complexity rather than removing it. The frontend inherits the backend’s invariants, so they have to be explicit, written down, and enforced by the fixtures.
  • An embedded integration is a cross-layer problem. Cookies, sessions, frame policies, and exits span the browser, the backend-for-frontend, and the identity platform; fix the layer that owns the cause, and be ready for the host’s answer to change late.
  • An observable product needs identity-stable telemetry. Otherwise the dashboard measures the instrumentation, not the users.
  • Strings shared across repositories — allowed origins, redirect destinations, contract versions — are a contract, and need an owner and a check.
  • A feature flag that gates a whole experience is an operational dependency. Plan its removal along with its introduction.
  • AI raises throughput only with an explicit verification layer: real contracts over cached conclusions, invariants written down, and nothing merged or posted without a person deciding to.