Engineering a data-heavy enterprise SaaS application
Context
The platform gives financial institutions risk scores for every asset in their portfolios, both current scores and projections. Its users are large banks, investment funds and portfolio managers.
I was with the client from 2020 to 2026, about four years of it on this platform, in a team of around three frontend and six backend engineers, working directly with the client's product team. We rebuilt the web application from scratch, because the one we inherited could not support the planned features or data volumes.
Role
As technical lead I owned the web application (React, Remix / React Router) and the Node BFF behind it. A separate team owned the Java services and the database. I was responsible for the stability of the application, planned work with the rest of the team, and built or integrated every planned feature. I helped choose the stack for the rebuild, and shaped how data flowed through the BFF, what was cached and where, and how the APIs were defined with the Java team.
Constraints
Portfolio size
At first a portfolio was limited to 2,000 assets, each with around 20 columns of scores. The API returned a whole portfolio in one response, so the browser held the full dataset in memory, and payloads grew with every asset added.
A full portfolio arrived as several megabytes of JSON. On low-end hardware the tab crashed. On newer machines it survived, with freezing that users noticed. That design set the ceiling, and the client needed much larger portfolios.
Tenant and user isolation
Each tenant could see only the data generated for its own portfolios. Some data, such as scores for publicly traded companies, was common to all tenants. Within a tenant, users had different roles, so data also had to stay separate between users.
Architecture
Data flow, before
Browser: React, AG Grid
│ holds the whole portfolio in memory
│ paging, sorting and filtering all happen here
▲
│ one response: the entire portfolio, several MB
│
Java backend, database
Every view a user asked for was served out of a dataset the browser already held, which is why the dataset had to be complete, and why the ceiling was where it was. Nothing was cached between requests, so the same datapoints were fetched again on the next load.
Data flow, after
I proposed this rework and carried it out: paging, sorting and filtering moved onto the Java backend, and Redis went in front of it to cut the number of calls reaching the backend and the database.
Browser: React, AG Grid
│ page, sort and filter are request parameters
▼
Remix / React Router
│
▼
BFF (Node): Zod validation, access checks
│ │
│ ▼
│ Redis
│ · score datapoints, keyed per user
│ · public company data, shared keys
▼
Java backend, database
paging, sorting and filtering run here
A grid view was not a single request and response. It was assembled from many datapoints, and the BFF did the assembly: it checked access, reused whatever was already cached, and fetched only the rest.
Caching
Score datapoints were cached in Redis with a TTL. As users moved around a portfolio, the BFF served the datapoints it had already seen from the cache and fetched only the missing ones from the database. That made exploring a portfolio faster and cut the number of calls needed to build each view.
API contracts
For each API, we and the Java team first agreed a spec in a shared document. I turned that spec into Zod schemas in the BFF, which validated every backend response at runtime. When the backend drifted from the spec, we found out in one of three ways: validation errors in the logs, failing tests, or the backend team telling us in advance.
Key decisions
From 2,000 to 20,000 assets
Raising the limit was not a fix in one layer. I proposed and carried out a redesign of the APIs and the data flow, and consequently of the whole web application, so that pagination, sorting and filtering happened on the server. The browser now held one page of results instead of the whole portfolio.
The limit went from 2,000 to 20,000 assets per portfolio, and the responses got smaller rather than larger: several megabytes became hundreds of kilobytes in the worst case. Ten times the data, a tenth of the payload. The crashes on low-end hardware stopped.
Testing strategy
- Playwright: end-to-end tests of user-facing features.
- Vitest and Testing Library: components, and the utilities that stitch data from several sources into one view.
Tradeoffs
We keyed cached data by user, not by tenant. Keying by tenant would have meant more cache hits and less memory, because users in the same organisation often look at the same portfolios. But users had different roles, and a single cache per tenant would have made a leak between users one mistake away. Per-user keys made that leak impossible by construction. We accepted the lower hit rate.
Public company data was the exception. It is the same for everyone, so it lived under shared keys and every tenant benefited from one cached copy.
Technical leadership and collaboration
I led the frontend and BFF side of the team for the six years I was with the client. Around three engineers looked to me day to day, and people from other teams did whenever a change crossed into the web application. What that meant in practice:
- Standards. Every feature that touched the web application or the BFF was built the way the team had agreed to build them: the stack, the Zod contracts, where data was cached, what was tested and how. Keeping that consistent across features was my job.
- Review. I was the required approver on that code.
- Onboarding. Three or four engineers joined the frontend side over those years. Getting them productive in a codebase this size was part of the role.
- Hiring. I screened and interviewed on every hire the team made.
- Planning. I did not run planning, but I was in it every cycle, because what the frontend could absorb depended on things only the frontend team knew.
Any change that touched the Java APIs, the BFF and the web application together went through three steps: a joint design session with the backend team, a written proposal, and sign-off from the client's product team.
Much of my pushback was about scope. Requests often involved more work than people expected, so I made the actual size of the work explicit before we committed to it. That was the part of the job that mattered most, and the part that looked least like engineering.
Outcome and lessons
Portfolios could grow ten times larger while the data crossing the wire got smaller, and the crashes stopped. Fewer backend calls and smaller payloads lowered what the platform cost to run. Dozens of features shipped on the new foundation over the following years, and I stayed with the client for six.
What I would do differently:
- Design for server-side data from day one. Returning everything at once was the reason for the rebuild.
- Treat cache invalidation as part of every write, not something to add later.
- Make scale expectations explicit early. Much of the pushback came from underestimating how large the workload was.