Corporate GenAI Platform Design Must Prioritize Requirements Before Technology Choice
September 30, 2026 · 3 min read
Also published in Polski, Português (Brasil), Norsk
A GenAI project often begins with a stack of technologies such as LangChain, vector databases, Kubernetes, Kafka, and a language model, but for enterprise platforms this early focus can be premature. The core architectural questions emerge not at the point of calling a large language model but around it, concerning data access, permission checks, independent scaling components, provider outage behavior, and the minimal set of features needed for the first version. Before selecting any stack, the platform's purpose, user base, functional, operational, and security requirements must be fixed. The key concept here is "authorized". Users must authenticate through a corporate identity provider, receive access only to permitted data, view answer sources, and have all interactions with protected data logged for audit. From these requirements arise architectural constraints that are not visible in a simple statement of building an enterprise RAG system. Non-functional requirements are easily confused with implementation choices; the system must survive component failures, but this does not mandate Kubernetes for the minimum viable product. Horizontal scaling can be planned without automatic scaling until real load appears. Architecture is defined not only by what the system does but also by the conditions under which it must operate.
In the C4 context, the outer system view clarifies who interacts with the platform and what lies behind its boundary. Three primary user types are identified: corporate identity provider, corporate document stores, and corporate SIEM or security monitoring system. The boundary of the system is defined by responsibility and lifecycle, not physical location; for example, a local LLM may run on a separate GPU cluster but remains an internal component if the inference infrastructure is part of the platform's C4 context, while the corporate identity provider, though physically nearby, is external to the platform's lifecycle. This perspective simplifies the subsequent container architecture. At the container level, the actual stack includes a React TypeScript web application and a Python document ingestion worker; while a target architecture may eventually contain a local LLM, Redis, reranking, and tool execution, the MVP does not require all these components. The critical decision is to avoid turning the Core API into dozens of microservices; it remains a modular monolith with logical domains such as Identity, Chats, Documents, Authorization, Audit, and Integrations, but these are not separated into independent deployable services early on. Extracting a service is justified only when a component exhibits distinct characteristics such as different load profiles, as with the Document Ingestion Worker which handles lengthy text extraction, chunking, and embeddings and must scale separately from the API.
The same principle applies to the AI Orchestrator, which has a different runtime environment and load pattern. This approach preserves modularity without incurring the early costs of distributed systems like network failures, distributed tracing, contract management, versioning, and complex deployment. While a full .NET implementation was possible, and a pure Python backend was also feasible, the choice to use .NET 10 for the application and corporate layer and Python only where AI/ML ecosystems are essential reflects an architectural trade‑off, not a language war. The additional complexity of maintaining two runtimes, dependency ecosystems, CI/CD pipelines, tracing, and contracts between services is accepted only where it delivers practical benefit, allowing AI workloads to scale independently while keeping business logic and security within the primary backend stack. A common follow up question is why not adopt a dedicated vector database immediately; specialized engines such as Qdrant or Weaviate appear natural for RAG, but introducing a separate vector store creates a second data repository and requires separate decisions about its management.