# CLAUDE Source: https://buildingfor.vc/CLAUDE # Building for Venture Capital - Content Guidelines ## Project Overview This is a technical book for engineers, data people, and technical operators building technology at VC funds. It's opinionated, practical, and based on real experience providing engineering services for venture capital and private equity through Boolean Industries, and at EQT and Inflection. ## Writing Voice and Tone **Direct and opinionated**: This book takes strong positions based on real experience. Use clear language: * ✅ "Don't build your own CRM. Use Attio or Affinity." * ❌ "You might want to consider whether building a CRM is the right choice for your organization." **Second-person "you"**: Always address the reader directly as someone doing this work. * ✅ "You're building for 5-30 people..." * ❌ "One might build for teams of 5-30 people..." * ❌ "Engineers build for teams of 5-30 people..." **Technical but accessible**: Assume technical competence but explain context. The reader is a competent engineer who may not know VC-specific details. **No corporate speak or buzzwords**: Avoid marketing language, hype, and vague business terminology. * ✅ "Use Postgres. It's boring technology that works." * ❌ "Leverage cutting-edge database solutions to maximize operational synergies." **Anti-hype**: Call out trends that don't matter. Be skeptical of new technology unless there's a clear practical benefit. ## Content Principles **Practical over theoretical**: Every section should answer "what should I actually do?" Include real examples, specific tools, actual code. **Acknowledge trade-offs**: Most decisions have context. Explain when advice applies and when it doesn't. * "For funds under 50 people, use \[X]. Larger funds might need \[Y]." * "This works if you're technical. If you're not, buy \[vendor] instead." **Real examples from experience**: Reference actual work through Boolean Industries, EQT, and Inflection when relevant. Be specific: * ✅ "At Inflection, we used Attio for CRM and synced it to Postgres nightly." * ❌ "Some funds use CRM systems and might integrate them with databases." **Strong opinions, weakly held**: Take clear positions but acknowledge alternatives and admit uncertainty when it exists. ## Structure and Formatting **Start sections with clear thesis statements**: First sentence should tell the reader what they'll learn or what position you're taking. **Use "The bottom line" summaries**: Many chapters end with a "Bottom Line" section that distills key takeaways. These should be practical, action-oriented summaries. **Author Notes in `` components**: Use for personal asides, contextual notes, or clarifications: ```mdx theme={null} **Author note**: At Inflection, we tried building this ourselves and it was a mistake. Save yourself the pain. ``` **Code examples with context**: Always include language tags. Explain what the code does and why, not just what it is: ```typescript theme={null} // TypeScript with Zod for validation import { z } from "zod" const CompanySchema = z.object({ name: z.string(), // ... rest of schema }) ``` **Real examples clearly labeled**: When providing case studies or specific examples: ```mdx theme={null} **Example: At Inflection, we used Mistral for PDF reading** When processing pitch decks... ``` ## Style Specifics **Sentence structure**: Prefer shorter sentences and clear paragraphs. Break up long blocks of text. Use periods over em-dashes in most cases. **Lists for clarity**: Use bulleted or numbered lists to make information scannable: * When listing tools or options * When outlining steps in a process * When comparing approaches **Avoid passive voice**: Write actively. * ✅ "Use Postgres for your data warehouse" * ❌ "Postgres should be used for data warehousing" **Be precise with technical terms**: Use correct terminology. Don't say "database" when you mean "data warehouse". Don't say "AI" when you mean "LLM". **Code over prose**: When showing how to do something, include actual code examples rather than just describing in prose. ## What This Book Is Not **Not comprehensive**: Cover what matters for VC funds. Skip edge cases that don't apply to the audience. **Not vendor-neutral**: Recommend specific tools that work. Don't try to cover every option. **Not future-focused**: Focus on what works today. Acknowledge emerging trends when relevant, but don't speculate about technology that doesn't exist yet. **Not academic**: This is practitioner knowledge. Cite specific experiences over research papers. ## Content Strategy **Evergreen when possible**: Focus on principles and approaches that will remain relevant as specific tools change. **Update for accuracy**: Keep tool recommendations current. If vendor landscape changes, update recommendations. **Real problems first**: Start from problems VC funds actually face, not from technology looking for applications. ## Technical Standards **Format**: MDX files with YAML frontmatter * title: Clear, descriptive (e.g., "Data Modeling and Schema Design") * description: Concise, practical summary (e.g., "How to model companies, deals, and relationships - the core data structures for VC infrastructure") **Code blocks**: Always include language tags for syntax highlighting ```typescript theme={null} // Good - has language tag ``` ``` // Bad - no language tag ``` **Internal links**: Use root-relative paths when linking between chapters: * ✅ `[Chapter 5](/part-2-tech-stack/crm-and-deal-flow)` * ❌ `[Chapter 5](../part-2-tech-stack/crm-and-deal-flow.mdx)` **External links**: Include for tools, vendors, and resources mentioned. Link to official documentation. ## Working With This Content **Push back when needed**: If an edit doesn't match the book's voice or would make the content less useful, say so and explain why. **Maintain consistency**: Check existing chapters for patterns in structure, terminology, and examples before adding new content. **Preserve strong opinions**: Don't soften the book's positions to make them more "diplomatic". The directness is a feature. **Keep it practical**: Every addition should help someone actually building VC infrastructure. If it doesn't, cut it. ## Git Workflow * Create feature branches for changes * Write clear commit messages * NEVER use `--no-verify` or skip pre-commit hooks * Commit frequently with logical groupings ## Do Not * Add fluff or filler content to hit word counts * Use buzzwords like "synergy", "leverage", "paradigm", "revolutionary" * Make recommendations without explaining why * Include untested code examples * Write in corporate or marketing voice * Soften opinions to be "safe" * Add features that don't solve real VC fund problems * Speculate about future technology developments * Make assumptions - ask for clarification when context is unclear # Contributors Source: https://buildingfor.vc/guide/contributors People who have helped shape this guide through feedback, contributions, and expertise. This guide wouldn't exist without the people who've contributed their time, expertise, and feedback. Thank you! * **[Ebba Forsberg](https://www.linkedin.com/in/ebba-forsberg-6a3820183/)** - Investor at [Norrsken Evolve](https://www.norrskenevolve.vc/) * **[Nik Niklaus](https://www.linkedin.com/in/nikolai-niklaus/)** - Formerly founded Whisper, a signal data platform for early stage VC funds (acquired by [Evertrace](https://www.evertrace.ai/)) * **[Ties Boukema](https://www.linkedin.com/in/tiesboukema/)** - Head of Data, Tech & AI at [Dawn Capital](https://dawncapital.com/) Additionally, thank you to everyone I've worked with at [Inflection](https://inflection.fund) and [EQT](https://eqtgroup.com/). The ideas in this guide came from building alongside you. *** Want to contribute? Check out the [GitHub repository](https://github.com/alexpatow/building-for-vc) to open an issue or submit a pull request. # Building for Venture Capital Source: https://buildingfor.vc/guide/index A technical guide to building technology for VC funds: from understanding the domain to shipping production systems. ## Welcome Most people building technology for VC funds are coming from the outside. They're talented engineers who haven't worked in venture capital before, don't yet know how funds actually operate, and often spend early cycles building the wrong things. This guide exists for the love of the game. I'm not trying to commercialize it or build a business around it. I just want to help funds think more holistically about technology and data in VC, and give engineers entering this space the resources I wish I'd had. **A note on perspective**: This guide reflects my personal experience building VC infrastructure. Your mileage may vary. Every fund is different, and what worked for me might not be the right approach for yours. Use this as a starting point, not gospel. I'm sure there are things I've missed. Dissenting opinions welcome! ## What This Guide Covers The guide is organized into three parts: Learn how venture capital funds actually work, understand your specific fund's needs, avoid common mistakes, and hire the right people for your data team. Explore the technology landscape at VC funds: what tools matter, when to build versus buy, and real examples from working funds. Deep dives into data providers, modeling, entity resolution, warehousing, integrations, security, and emerging trends like MCP and AI agents. ## Who This Is For **You should read this if you:** * Just joined a VC fund as an engineer or CTO * Are building tools or infrastructure for venture capital * Want to understand how VC funds operate from a technical perspective * Need practical guidance on data modeling, integrations, and architecture for VC **This guide is not:** * An introduction to venture capital investing * A guide to becoming a VC or raising money * Generic startup or tech company advice ## Getting Started Begin with the fundamentals of how VC funds work - essential context before building anything. Skip to specific topics or use this as a reference guide if you're already familiar with VC basics. ## About the Author [Alex Patow](https://www.linkedin.com/in/alexpatow/) runs [Boolean Industries](https://boolean.industries), providing engineering services for venture capital and private equity. He's been building in VC since 2020, including at [EQT](https://eqtgroup.com/) (Motherbrain) and [Inflection](https://inflection.fund/). He's been recognized in the [Data Driven VC Landscape](https://datadrivenvc.io/reports) as a leader in the field and has given multiple talks on the subject. ## Contributing Found an error, have a suggestion, or want to contribute additional content? Contributions are more than welcome! Check out the [GitHub repository](https://github.com/alexpatow/building-for-vc) to open an issue or submit a pull request. See the [contributors](/guide/contributors) who've helped shape this guide. ## Disclaimer The views and opinions expressed in this guide are solely my own and do not necessarily reflect the official policy or position of any employer, client, or organization I am or have been affiliated with. Any product or vendor recommendations are based on personal experience and are not endorsements by any affiliated entity. # Common Mistakes When Starting at a Fund Source: https://buildingfor.vc/guide/part-1-understanding-vc/common-mistakes Learn the seven common mistakes developers make when joining VC funds and how to avoid them before they cost you months of work. ## Overview You've learned the VC fundamentals. You've analyzed your specific fund. Now you're ready to start building, right? Not quite. Even with solid understanding of VC and your fund's needs, there are common mistakes that almost every technical person makes when they join a VC fund. I've made most of them. Other CTOs at funds have made them. Developers building VC software have made them. This chapter exists so you don't have to learn these lessons the expensive way. ## Mistake #1: Jumping In Too Early Picture this: It's your first week at a new fund. You walk in and immediately see problems everywhere. Excel spreadsheets tracking deals. Manual processes for everything. No proper CRM. Partners complaining about how hard it is to find information. You're a builder. You see technical problems. You want to fix them. So you start coding. You're excited. You're making progress. You're shipping features. Three months later, you demo your beautiful new deal flow tool to the partners. They're polite. They say it's nice. But nobody uses it. Or worse, they try it for a week and go back to their spreadsheets. You built the wrong thing. This happens because you're hired to "build technology," so you feel pressure to ship quickly. Technical problems are comfortable. Domain problems are messy and ambiguous. Building feels productive. Research feels slow. You want to prove your value fast, and the best way you know how is to write code. But here's the reality: the first one to three months should be mostly research and observation, not building. You need to understand how the fund actually operates, not how they say they operate in the interview. You need to understand what the real pain points are, not what people think they are when you ask them directly. You need to know what's been tried before and why it failed. Most importantly, you need to understand what workflows are essential versus nice-to-have. **Author Note: Jumping in too early at Inflection** At Inflection, we started by building a sourcing platform including agents, multiple data sources, and a graph database for modeling relationships. It showed some promising results, but didn't deliver the value we were hoping for. The interesting companies surfaced by the tool? We were already talking to them. Growing the haystack didn't help us find better needles. ### How to Avoid It Spend your first two weeks just observing. Shadow each GP for a day. Attend all the partner meetings. Watch how deals actually flow through the organization. Take notes. Ask questions. But don't propose solutions yet. You don't know enough. **Author Note: Shadowing GPs led to Kepler** [Kepler](https://svrgn.substack.com/p/introducing-kepler-inflections-home), Inflection's research platform, wouldn't have come about if it weren't for shadowing the GPs in a few deals and realizing how much research we were doing on the companies. That research wasn't being stored in a great way, it was too siloed. So we decided to build a research platform. The insight came from observation, not from asking "what should I build?" Weeks three and four, start documenting what you've seen. Map out the current workflows, not the ones they told you about in the interview, but the ones you actually observed. Identify pain points. Document the current tools and data sources they're using. Interview each team member individually. You'll be surprised how different the story is when you talk to people one-on-one versus in a group setting. Weeks five through eight, prove you understand the domain before building anything big. Fix one small, obvious problem. Automate one manual process that everyone agrees is annoying. Show that you get it. Build trust before building big. **Author Note: Starting small at Inflection** At Inflection, I dipped my toes in by helping to automate the process of creating valuation memos for our audit process. This was a big win for the ops team which had to compile these manually, and it let us test out some AI frameworks for future tasks. Small win, real value, built trust. By week nine, you should have enough context to propose a real technical strategy. Get buy-in on priorities. Now you can start building the right things. Watch out for these red flag phrases, especially from yourself: "Let's just start building and iterate." "This should be quick to throw together." "We can always change it later." These are all signs you're jumping in too early. There's one exception: if the fund is two people and has literally no technology at all, you may need to move faster. But even then, start small and specific. Build one thing that solves one problem. Then build the next thing. ## Mistake #2: Not Aligning on Buy vs. Build Strategy Here's a nightmare scenario: You spend three months building a custom deal flow CRM. You're proud of it. It's tailored perfectly to the fund's workflow. Then in a partner meeting, the managing partner casually mentions they're also evaluating Affinity and Attio. Wait, what? Nobody told you they were considering buying software. Or the reverse happens. You do your homework, research the available tools, and propose using Affinity for deal flow management. The GPs look confused. They hired a "CTO." They expected you to build something custom. They're visibly disappointed that you want to buy instead of build. This misalignment happens all the time. The role of "CTO at a VC fund" means different things to different people. Some funds want someone to build custom software. Others want someone to evaluate, buy, and integrate existing tools. Most want some combination, but nobody makes that explicit upfront. Add to this that fund partners often don't know what's available to buy in the VC software ecosystem (it's niche and not well marketed), and you don't know either when you first start. Sometimes there's internal politics: some partners want to build custom tools, others just want to buy and move on. So when should you buy versus build? Commodity functionality should almost always be bought. CRM basics, email tracking, document signing. These are solved problems. Complex compliance requirements like fund administration and LP reporting should definitely be bought unless you have deep expertise and significant resources. Standard integrations with banks, DocuSign, and other services work better when you're using established tools. If you're a small team without significant engineering resources, buy more and build less. And if speed to value matters more than perfect fit, buying gets you there faster. Build when you have true differentiation. If your fund has a novel sourcing strategy or unique thesis that requires custom tooling, that's worth building. Build when you have highly specific workflows that genuinely don't map to existing tools (but be honest about whether they're truly unique). Build when you have the engineering resources to maintain what you create. Build when you've actually tried existing tools and they've failed for specific, articulable reasons. There's also a middle ground: build on top of bought software. Use vendor APIs to extend functionality. Build internal tools that integrate with external platforms. Create custom reporting and analytics on top of existing data. This often gives you the best of both worlds. The key is to have this conversation explicitly, not assume everyone's on the same page. In the interview process, ask directly: "What do you expect me to build versus buy?" Ask what tools they've already evaluated or tried. Ask about their budget for software versus engineering salaries. In your first month, create a landscape document of available VC tools. Categorize them: deal flow, portfolio management, LP reporting, fund administration. Propose a "buy/build/integrate" strategy and get explicit sign-off on the approach. Have the hard conversation early: "Here's what exists that we could buy. Here's what we'd need to build. Here's the tradeoffs in terms of time, cost, and fit. What's more important to you: speed or perfect fit?" Get this alignment before you write a single line of code or sign a single software contract. **Author Note: What not to build anymore** I've seen funds try to build custom CRMs, which used to be meaningful differentiation. But now you can get most features off the shelf with tools like Affinity or Attio. As a solo engineer or small team, I wouldn't spend time building a new one from scratch. The same goes for data scraping. I've seen engineers spend weeks building LinkedIn scrapers. While that data is important, your time as an engineer is the most expensive resource. Don't waste it when you can buy data from providers like [People Data Labs](https://peopledatalabs.com/), [Crunchbase](https://www.crunchbase.com/), or others. Build on top of purchased data, don't recreate it. ## Mistake #3: Not Having Enough Resources You're the only technical person at a five-person fund managing \$50M. In your first week, you collect all the requests. They want a custom deal flow CRM. Portfolio tracking dashboards. Automated LP reporting. A public-facing fund website. Internal knowledge management. Data pipelines from five different sources. You write it all down. You start estimating. This is two to three years of work. Minimum. For one person. This happens because funds fundamentally underestimate software complexity. They don't understand what goes into building and maintaining custom software. And honestly, you probably overestimate what you can ship alone too. You're optimistic. You're capable. You think you can move fast. Nobody scoped the work before hiring you. The interview focused on vision and potential, not realistic timelines. The "how hard can it be?" mentality prevails. Here's what one technical person can realistically do: integrate and manage existing tools, build one or two custom internal tools (simple ones), automate some workflows, create custom reports and dashboards, and manage data infrastructure. That's a full plate. That's valuable work. Here's what one technical person cannot do: build and maintain a full custom stack from scratch, replace Carta, Affinity, and DocuSign with custom-built alternatives, provide 24/7 support for critical systems, or keep up with every new request that comes in. You need to set expectations early. During the interview process, ask to see the list of desired projects. Ask about team size and budget for contractors or vendors. Be honest about what's realistic. Discuss priorities: what actually comes first? In your first month, audit all the requested projects. Estimate time for each one generously (double your initial estimate, then add 50%). Show the math: "This is three years of work. I'm one person. Let's talk about priorities." Set clear expectations about what you can ship in Q1, what requires buying software, what requires hiring another person, and what you probably shouldn't do at all. Get help where you need it. Budget for contractors for specific projects. Use agencies for one-off work like website design. Automate with no-code tools where possible. Buy software for non-differentiating work. Watch out for red flags: "We need custom everything" usually means nobody's thought about buy versus build. "Buying software is expensive" ignores that building is far more expensive when you factor in your time. "You'll have help... eventually" is code for "we have no concrete plans to hire." "Just get something working, we'll improve it later" is how you accumulate crushing technical debt. **Author Note: The evolving backlog at Inflection** When we started building at Inflection, we had this huge backlog of products we thought we needed (you can [find a lot of them in this list](https://svrgn.substack.com/p/an-engineering-approach-to-venture) from before I was hired). If I was to execute on all these projects, I would always be context switching and stretched too thin. The key was constantly setting realistic priorities with the partnership, spelling out what it would actually take to build things, and forcing them to give feedback on what mattered most. Over time, many of those "critical" needs changed. The fund evolved. More importantly, more products became available to buy off the shelf. Half the backlog became irrelevant or solvable with purchased tools. Keep re-prioritizing based on what's actually needed today, not what seemed important six months ago. ## Mistake #4: Over-Engineering for Scale You Don't Have You're building a deal flow tool for a fund that sees 200 deals per year. You sit down to design the architecture. Microservices, obviously. Event-driven workflows for flexibility. Caching layers for performance. Kubernetes for orchestration. Horizontal scaling for growth. This is "best practice," after all. This is how you build software "the right way." Six months later, you've spent all your time on infrastructure. The tool still doesn't actually work for the core use case. The fund is frustrated. Partners are asking why this is taking so long. You're debugging Kubernetes networking issues instead of shipping features. This happens constantly when engineers come from tech companies with real scale problems. You're used to building for millions of users. You want to build "the right way" from the start. Over-engineering feels professional and impressive. Simple solutions feel almost embarrassingly basic. Sometimes it's even resume-driven development: you want to list these technologies on your LinkedIn. But here's the reality. A \$100M VC fund's scale looks like this: five to twenty employees, 200 to 500 deals per year in the funnel, 20 to 50 portfolio companies, 20 to 100 LPs, and quarterly reporting (not real-time dashboards). This is not scale. This is a small business. You can run all of this on a single Postgres database, a monolithic application, simple cron jobs, and basic authentication. You don't need microservices. You don't need Kubernetes. You don't need event-driven architecture. You don't need complex caching. You don't need auto-scaling. What you need is to ship something that works. Start with the simplest thing that could possibly work. Single application server. One database. Cron jobs for batch processing. Deploy to Vercel, Render, or Railway with a single command. Manual processes for rare operations that happen once a quarter. Add complexity only when you have a specific problem. The database is slow? Add indexes and optimize queries. Deployments are risky? Add tests and CI/CD. The server goes down? Add monitoring and auto-restart. But don't build complex caching because something "might be slow someday." Build it when it's actually slow today. Know your actual constraints. How many users? Probably fewer than 20. How much data? Probably less than a terabyte. How many requests per minute? Probably fewer than 100. What's acceptable downtime? Probably an hour is totally fine. This is an internal tool, not a public API. You'll know you need more complexity when you have actual performance problems (not hypothetical ones), when the simple solution is causing real pain (not theoretical pain), when you've exhausted simple optimizations, and when the ROI of additional complexity is crystal clear. **Author Note: Kubernetes at a VC fund** Our initial sourcing tool at Inflection was set up on Kubernetes for workflow management, using Terraform. All the sins I told myself I wouldn't commit. Over a week, we transitioned to [Modal](https://modal.com/). We ended up saving ourselves a bunch of money on resources we didn't need and had a much more reliable tool. Sometimes you need to catch yourself mid-over-engineering and course correct. ## Mistake #5: Not Understanding Data Sensitivity and Compliance VCs have confidentiality obligations to LPs and portfolio companies. What does that mean for your technical decisions? You build a cool deal flow tool. You add a public API so you can integrate with other services. You set up Slack notifications for new deals. You log everything to Datadog for monitoring. You use OpenAI's API to automatically analyze companies and generate summaries. It all works great. Modern development practices. Best in class tools. Then someone points out you've exposed confidential deal information, LP identities, and fund strategy to multiple external services. Disaster. A year later, during an LP audit, auditors ask for historical valuations with complete audit trails. You can't produce it. Your system overwrites old values. The auditors are not happy. VC confidentiality and compliance requirements are fundamentally different from consumer apps. In consumer tech, data is often public or semi-public. You move fast and break things. But in VC, there are real confidentiality obligations and audit requirements. The challenge is finding the right balance between security and actually building useful things. VCs tend to care about this less than PE funds, where everything is top secret. VCs frequently need to syndicate deals with other investors, so some information sharing is necessary. Different firms have different risk tolerances. Some are more open, others more locked down. Your job isn't to make those decisions. It's to understand where your partnership stands on that spectrum and build accordingly. Here's what's typically highly sensitive: LP identities and commitments (never public, often contractually confidential), competitive deal flow, detailed diligence notes (legal liability if they leak), fund performance details (LPs or potential LPs only), portfolio company board materials, and specific investment terms (covered by NDAs). You have legal obligations through NDAs, LP agreements, and securities regulations. VC funds also get audited. Annual audits by accounting firms. LP due diligence before commitments. Sometimes regulatory audits. Auditors need historical data, audit trails, source documentation, and process documentation. This isn't optional, but the level of rigor varies by fund size and structure. In your first month, have explicit conversations: "What data can never leave our systems?" "What are our confidentiality obligations?" "What's our risk tolerance around third-party tools?" "What audit requirements do we have?" Review the actual LP agreements. Understand the constraints, but also understand where there's flexibility. The trap is being so paranoid about confidentiality that you can't build anything useful. No third-party tools means no productivity. No AI means no automation. No integrations means manual processes everywhere. That's not the answer either. Find the balance your partnership is comfortable with. Maybe you can use foundational model APIs for non-sensitive tasks but not for deal flow analysis. Maybe you can use cloud logging with proper data filtering. Maybe you use third-party tools but with specific data exclusions. The key is understanding the risks you're putting the firm at, getting explicit approval for your approach, and documenting the decisions. Build for both confidentiality and auditability where it matters. Role-based access control. Audit logs. Data encryption at rest. Version history on financial data. Soft deletes, not hard deletes. But don't go overboard on things that don't matter. Document your processes for the things that need documentation. It's not your call what gets shared with whom. But it is your responsibility to understand the boundaries, propose solutions that work within them, and make the tradeoffs explicit. **Author Note: Work with compliance, don't fight it** At a small fund, there's an advantage: you often *are* compliance. You get to make the rules (and own the outcomes). At larger funds (or funds that are part of larger organizations), you need to build strong relationships with the CISO and internal IT teams. If you're running fast and want to try the latest tools, building trust with compliance is critical to moving quickly. Figure out what's important to them. Pick your battles. But most importantly, collaborate with them rather than maintaining the usual "devs vs. compliance" battle that persists at most companies. They're usually great people who want to help, not block you. ## Mistake #6: Building in a Vacuum You spend three months building a beautiful portfolio dashboard. You've thought of everything. Custom visualizations. Drill-down capabilities. Export features. It's elegant. It's fast. It's well-architected. You're proud. You demo it to the partners. They look at it. "This is nice," one says. "But it's not what we need." They wanted something completely different. You built in isolation without regular feedback. Three months of work that missed the mark. You wanted to surprise them with a finished product. You were embarrassed to show ugly work-in-progress. You thought you understood the requirements from those initial conversations. The partners are busy with deals and you didn't want to bother them with constant check-ins. You're used to longer release cycles from your previous job where you'd ship quarterly. But internal tools need constant feedback. Partners are experts at investing, not at articulating product requirements. They think in terms of outcomes and workflows, not features and data models. Requirements evolve as they start using early versions and realize what actually matters in practice. Workflows evolve as the fund grows. Your initial understanding is always incomplete, no matter how many questions you asked. This isn't a failure on anyone's part. It's just how product development works when you're solving complex, nuanced problems. The cost of building wrong is severe. Three months of wasted time. Loss of trust (now they doubt you understand their needs). Missed opportunity cost (you could have been building the right thing). Technical debt from choosing the wrong architecture for the wrong problem. Ship incrementally. Build the simplest possible version first. Get it in their hands. Learn what's wrong. Iterate based on real usage, not hypothetical requirements. Create feedback loops. Weekly demos to the team. A Slack channel for feature requests. Regular one-on-ones with heavy users. Usage analytics showing what features actually get used. Embrace being wrong: "I built this based on what I thought you needed. What's wrong with it?" "Here are three options. Which direction feels right?" "This is a prototype to test the concept." Show progress regularly. How often depends on your team size and what you're building - but err on the side of more frequent, shorter check-ins rather than big quarterly reveals. **Author Note: Embedding with the business** A great model I saw was having someone from the investment team act as a part-time "product manager" during major rollouts. They understood the workflows, could prioritize features, and gave real-time feedback. It's a big ask, but even a few hours a week from the right person makes a huge difference. ## Mistake #7: Treating Investors Like Engineers You're in a partner meeting presenting your solution. "We'll use a denormalized data model with materialized views for performance," you explain. The partners nod politely. The conversation moves on quickly. You realize they're not engaging with the technical details because those aren't the details that matter to them. Or you sit down with a partner expecting them to write detailed requirements with acceptance criteria, like product managers do. They describe the outcome they want: "I want better deal tracking." You ask for more specifics about fields and workflows. The conversation stalls. You're speaking different languages. You're used to working with other engineers where technical explanations are how you communicate. You expect someone to write specs and acceptance criteria because that's how product development works at tech companies. But investors are experts at evaluating companies and making investment decisions. Their expertise is in understanding markets, founders, and business models, not in articulating software requirements. Investors think in outcomes, not implementations. "I want to see which deals are getting stale." "I need to prepare for quarterly LP calls." "I can't remember who introduced this company." The value is in solving the problem, not in how you solve it technically. They trust you to make the technical decisions. Translate technical decisions into business value. Don't say "I'll implement real-time sync with a WebSocket connection." Say "The data will update immediately when anyone makes changes." Show, don't tell. Don't explain the architecture. Show a working prototype. Let them experience the solution. Ask outcome-focused questions. Not "What fields should be in the Deal model?" Ask "What do you need to know about a deal to make a decision?" Their requirements come as business needs, not technical specifications. Extract requirements through conversation and observation. Build something, show it, refine it based on how they actually use it. Be the translator. They describe the business problem: "I want better deal tracking." You translate that into technical requirements: deal stage visibility, activity history, and reminders. You build that. You show it. You adjust based on their feedback. This is your expertise. This is why they hired you. There are exceptions. With the CFO or fund administrator, you can get technical. They understand data, processes, and compliance. Technical discussions are productive. They can help specify requirements. With engineers you hire later, technical depth is expected. Implementation discussions are valuable. Architecture reviews are helpful. But with GPs? Focus on outcomes. **Author Note: Wear the product hat** This has been a big learning for me being the only engineer at Inflection: you need to wear the product hat, probably more than the engineering hat. Be really great at showing new features, but also explaining when and how to use them. Write documentation. Set metrics and follow up on those metrics. Kill what isn't working. You own the outcomes, not just the code. That's the real job. ## The Bottom Line These seven mistakes are common and understandable. Every technical person joining a VC fund faces some version of these challenges. The difference between success and frustration isn't talent or experience, it's knowing the mistakes exist and planning accordingly. Before you write any code, spend meaningful time understanding the fund. Align explicitly on buy versus build strategy. Be realistic about what you can accomplish with available resources. Start simple and add complexity only when you have specific problems to solve. Understand data sensitivity and confidentiality from day one. Build for auditability, not just functionality. Get feedback early and often. Communicate in outcomes, not technical implementations. Don't jump straight into building. Don't assume you know what's needed based on interviews alone. Don't over-engineer for hypothetical scale. Don't expose confidential data through modern development practices. Don't ignore compliance requirements until audit season. Don't build in a vacuum for months without feedback. Don't expect detailed specifications from non-technical partners. Don't use technical jargon when discussing solutions with GPs. The mistakes aren't failures. They're learning opportunities if you catch them early. Most developers make at least half of these mistakes in their first year. The goal isn't perfection. The goal is awareness and course correction. You've now learned the fundamentals of how VC works, how to analyze your specific fund, and the common mistakes to avoid. Next, we'll cover [hiring your data team](/guide/part-1-understanding-vc/hiring-your-data-team): when to hire, what to look for, and how the role evolves. # Hiring Your Data Team Source: https://buildingfor.vc/guide/part-1-understanding-vc/hiring-your-data-team When to hire your first data person, what to look for, and how the team evolves. ## Overview The first technical hire at a VC fund is a unique role: part engineer, part product manager, part internal consultant. Getting it right means finding someone who can operate independently in an ambiguous environment while delivering real value. This chapter covers when to hire, what to look for, and how to grow the team over time. ## When to Hire Your First Data Person The best funds hire proactively, before they realize they need it. But the right timing depends heavily on fund size and strategy. **Signals that you're ready:** * Partners are spending significant time on manual data tasks that could be automated * You have a specific technical initiative that needs leadership (not just a vague "we should be more data-driven") * The fund has conviction that data and technology will be core to differentiation, not just operational efficiency * You have budget for the hire and the tools, infrastructure, and data they'll need **Signals that you're not ready:** * "Data-driven" is a buzzword in your marketing, not a real strategy * You expect one person to solve all technical problems while you figure out what you actually need * There's no budget for data subscriptions, cloud infrastructure, or tooling beyond salary * Partners aren't willing to invest time in defining problems and giving feedback Many successful funds operate without dedicated technical staff. They use off-the-shelf tools like Affinity or Attio, outsource specific projects, and focus their energy on deal sourcing and portfolio support. The decision to hire should be driven by your fund's specific needs and strategy, not by what other funds are doing. But if you've decided to hire, here's what to look for. ## What to Look for in Your First Hire **Hiring advice from [Ties Boukema](https://www.linkedin.com/in/tiesboukema/) at Dawn Capital** The first tech hire should be able to sell the vision of what’s possible with data & tech internally and generate excitement among partners. They need to be self-motivated to identify opportunities where data, tech & AI can add value and pitch their own projects. Investors typically haven't built data pipelines & complex solutions that scale, so their intuition on what's possible (and how fast!) can be off. Great tech hires need to be proactive problem-identifiers who can independently scope and drive their own work rather than wait for assignments. This is different from typical banking or consulting analysts who are accustomed to working on urgent, well-scoped deliverables. You also need to be comfortable figuring things out largely on your own. If you come from a tech firm, there's lots of engineers to ask for help. You will almost certainly be the most technical person at the firm, which is a different environment. On the flip side, you get a level of autonomy (and speed!) that is extremely hard to find unless you're a founding engineer at a startup. How much you weight vision-selling versus hands-on execution depends on fund size and team structure. At larger funds where you're building a team, you can separate leadership (selling the vision, setting strategy) from execution (building the infrastructure). At smaller funds where one person does everything, vision-selling and technical depth have to coexist in the same hire. **The ideal technical profile:** * Broad software engineering skills * Data engineering depth (pipelines, databases, integrations) * VC domain knowledge Technical skills are learnable. Product sense in a specialized domain like VC is harder to acquire. **Hiring too junior is the biggest mistake funds make.** Someone early in their career, even if technically capable, often struggles without external structure. Look for someone who has shipped projects end-to-end with minimal supervision. ## How Hiring Evolves Don't hire a second person until the first hire has identified clear workstreams to divide. The first hire's job isn't just building tools. It's discovering what the fund actually needs: experimenting with different projects, identifying where sustained investment makes sense versus where quick wins are sufficient. **When to hire a second person:** * You have clear, separable workstreams (e.g., data infrastructure vs. internal tools vs. ML/AI initiatives) * The first hire is consistently underwater despite good prioritization * You've identified a big swing that requires dedicated focus and the first hire can't abandon other responsibilities **When hiring specialists (ML engineers, data engineers, etc.):** * Only hire specialists when there's enough work to keep them occupied long-term * If you want to take a big swing on AI/ML, use the first hire to scope it and build the data infrastructure to support it, then hire someone dedicated to that initiative * For niche specialties with uncertain long-term demand, use consultants for the initial build. Larger funds often do this: bring in external help to ship something, then decide whether the workstream justifies a full-time hire The pattern is: generalist first to figure out what matters, then specialists for sustained investment in specific areas. ## Compensation and Role Definition VC fund technical roles sit awkwardly in the market. You're not a big tech company with standard levels and compensation bands. You're not a startup with equity upside. You're a small team at a financial firm, which means different expectations on both sides. **On compensation:** * Base salaries are typically in line with tech companies * Carried interest (carry) is still rare for technical hires but becoming more common, especially at senior levels * Equity is non-existent (VC funds don't have equity like startups do) * The value proposition is: interesting problems, high autonomy, direct impact, exposure to the startup ecosystem **On role definition:** * Titles vary by firm. Some use "CTO," others prefer "Head of Data" or "Head of Engineering" * Be clear about scope: are they building tools for deal flow, portfolio analytics, LP reporting, all of the above? * Define the reporting structure: do they report to a partner? The COO? Operating independently? * Beyond formal reporting, identify which stakeholders outside the investment team they'll work with – operations, finance, investor relations – and ensure those people are bought in on the hire ## Red Flags in Candidates Watch out for these warning signs: **No evidence of independent ownership.** Ask about projects they drove end-to-end. If every example involves working within a well-defined scope set by someone else, they'll struggle in a VC environment. **Dismissive of "boring" work.** Data pipelines, integrations, and internal tools aren't glamorous. If a candidate only wants to work on interesting technical problems, they won't do the unglamorous work that actually matters. **Can't explain technical concepts simply.** They'll be working with non-technical people every day. If they can't translate technical trade-offs into business terms, they'll struggle to get buy-in for projects. (See [Mistake #7](/guide/part-1-understanding-vc/common-mistakes#mistake-%237%3A-treating-investors-like-engineers)). ## The Bottom Line Your first hire needs broad software skills with depth in data engineering and VC domain knowledge. They should be able to sell the vision and deliver on it. Avoid hiring too junior. Let your first hire identify what actually matters before expanding the team. In Part 2, we'll shift from understanding VC to the practical work of building: what tools exist, when to buy versus build, and how to put together a tech stack that serves your fund's needs. # Understanding Your VC Fund Source: https://buildingfor.vc/guide/part-1-understanding-vc/understanding-your-vc-fund Learn how to analyze your specific fund's stage, strategy, and workflows to build the right technical solutions. ## Overview Now that you understand the fundamentals of how VC works, it's time to understand the specific fund you're building for. This is where most developers go wrong. They build generic "VC software" without understanding that a pre-seed fund needs completely different tools than a growth equity fund. The key insight: Not all venture capital is the same. Building software for a pre-seed fund is fundamentally different than building for a growth equity fund or a PE firm. If you don't understand your fund's stage, strategy, and workflow first, you'll build the wrong thing. This chapter will teach you how to analyze your fund and translate that understanding into technical decisions. ## The Critical First Question Before writing a single line of code, you need to answer: **What type of fund am I building for?** This isn't academic. The stage focus, investment strategy, and fund structure determine what data you need to track, how complex your workflows need to be, what integrations matter, how much automation is possible, and what your data model should look like. Building generic "VC software" is a recipe for failure. You need to build for a specific fund type. ## Fund Types by Stage The stage a fund invests at fundamentally shapes everything about how they operate. A seed fund seeing 100 companies per month operates nothing like a growth equity fund doing 3-month diligence processes on 10 companies per year. ### Pre-Seed and Seed Funds (\$10M-\$100M) Seed funds vary widely in their approach, some are concentrated portfolios, other are more "spray-and-pray". A typical seed fund might make 5-10 investments per year with check sizes around \$500K-\$2M. These aren't tiny angel checks, but meaningful seed investments where the fund takes a significant ownership stake. These funds are often run by small teams of 2-5 people. There's no large analyst pool. Partners wear multiple hats: sourcing, evaluating, supporting portfolio companies, managing LP relationships. The diligence process is relatively quick (1-3 weeks typically) but still substantive. They're evaluating founder quality, market opportunity, early traction, and whether the company has a real shot at reaching Series A. Deal flow matters, but it's not about processing hundreds of deals. It's about seeing enough quality opportunities to make 5-10 great bets per year. The fund might review 200-500 companies annually to find those investments. Research and thesis-driven sourcing becomes important. Many seed funds actively map ecosystems and reach out to interesting founders rather than waiting for inbound. There's a subset of seed funds that take a different approach: 30-50+ investments per fund with smaller checks (\$50K-\$500K). These funds operate more like volume businesses with minimal diligence and binary outcomes. But this is less common than the concentrated approach. Examples: * [Inflection](https://inflection.fund/) - Deep tech pre-seed/seed * [Notation Capital](https://notation.vc/) - Concentrated seed fund * [Precursor Ventures](https://precursorvc.com/) - Pre-seed and seed * [Y Combinator](https://www.ycombinator.com/) - Accelerator with high-volume seed investing * [Tiny VC](https://tiny.vc/) - High-volume portfolio approach ### Series A and B Funds (\$50M-\$300M) Series A and B funds occupy the middle ground. They see moderate deal volume (20-40 companies per fund), write checks in the \$2M-\$10M range, and have substantial diligence processes. These funds usually take board seats. They reserve 30-50% of their capital for follow-on investments in their best companies. Deal teams are larger here. You might have 5-10 investment professionals, plus operating partners or specialists. The diligence process involves multiple people: someone does market research, someone handles financial modeling, someone does customer reference calls, maybe someone does technical evaluation. Everything feeds into an investment committee meeting where partners debate the deal. Examples: * [Earlybird](https://earlybird.com/) - European Series A/B * [Balderton Capital](https://www.balderton.com/) - European Series A/B * [Index Ventures](https://www.indexventures.com/) - Series A and B focus ### Growth Equity (\$500M-\$2B+) Growth equity funds operate at the other end of the spectrum. Low deal volume (maybe 10-20 companies per fund), very large checks (\$25M-\$100M+), and diligence processes that can stretch 3-6 months. These deals look more like acquisitions than venture bets. The fund is evaluating mature businesses with real revenue, established customers, and clear unit economics. The focus here shifts from "will this work?" to "can this scale profitably?" Customer references matter a lot. Financial models need to be detailed. The fund often co-invests with other growth funds, so syndicate management becomes important. There's significant operational involvement. Some growth funds effectively function as operating partners, helping professionalize finance, sales, and operations. Examples: * [Insight Partners](https://www.insightpartners.com/) - Global growth equity * [General Atlantic](https://www.generalatlantic.com/) - Growth stage investor * [Summit Partners](https://www.summitpartners.com/) - Growth equity and buyout ### Private Equity / Buyout Funds PE funds are a different world. Very low deal volume (5-15 companies per fund), controlling stakes or full acquisitions, complex capital structures involving debt, and extensive operational involvement. These aren't minority investments hoping for growth. These are acquisitions where the fund owns and operates the business. Deal sourcing happens differently. Deals often come from investment banks running formal processes, not warm intros from founders. Due diligence involves consultants, legal teams, accounting firms, and operational experts. The fund isn't just evaluating the business, they're planning how to grow revenue, improve operations, cut costs, and eventually sell it. Post-investment, the fund isn't just monitoring. They're actively managing. They might replace the CEO, restructure operations, or make add-on acquisitions to roll up competitors. Portfolio company management requires deep integration with the company's systems. Here's where the role of technology and data changes fundamentally. In PE, the questions deal teams ask are often so bespoke that the technologist's role shifts from building systems to being part of the deal team itself. You're using data to help deal teams get to conviction. Can we model the impact of operational improvements? What does the competitive landscape look like quantitatively? How do we assess add-on acquisition targets? Post-deal, the work continues with portfolio companies. You might help a portfolio company with M\&A analysis for add-on acquisitions. Or build custom models for their specific business challenges. Each project is tailored to bespoke needs. This isn't about building one platform that serves all portfolio companies. It's about being a data and technology partner embedded in the operational work of the fund and its portfolio. Examples: * [EQT](https://eqtgroup.com/) - Global PE with sector focus * [KKR](https://www.kkr.com/) - Large-cap buyout and growth * [Blackstone](https://www.blackstone.com/) - Largest alternative asset manager ### Multi-Strategy Funds Some of the largest and most prominent VC firms operate across multiple stages and strategies simultaneously. These multi-strategy funds run separate vehicles for different stages: a seed fund, a Series A/B fund, a growth fund, maybe even a crypto fund or bio fund. Each vehicle operates with its own strategy, team, and capital base, but under one brand and shared infrastructure. The structure creates interesting dynamics. A company might start as a seed investment from the early-stage fund, then receive follow-on investment from the growth fund years later. Deal flow can be shared across funds. A company that's too late for the seed fund might be perfect for Series A. Partners often specialize by stage or sector but can collaborate across funds. This multi-fund structure affects everything about operations. You're not building for one fund with one strategy. You're building infrastructure that needs to support multiple funds with different sourcing methods, evaluation criteria, workflows, LPs, reporting requirements, and team structures. But you also want shared infrastructure where it makes sense: one deal flow system that all funds use, one portfolio tracking platform, shared research and data. The challenge is managing complexity while maintaining the benefits of integration. Each fund needs its own LP reporting, its own capital tracking, its own performance metrics. But deals, companies, and relationships should be shared. A partner working on both the seed and growth funds needs visibility across both, but LPs for each fund should only see their specific fund's data. Examples: * [Andreessen Horowitz](https://a16z.com/) - Multi-stage from seed to growth * [Sequoia Capital](https://www.sequoiacap.com/) - Seed, venture, and growth funds * [Accel](https://www.accel.com/) - Early stage through growth ## Fund Investment Strategies Stage isn't everything. Two Series A funds can need completely different technical architectures based on their strategy. Strategy determines your data model, integration requirements, and workflow complexity just as much as stage does. Geographic focus, industry specialization, and investment approach all create different technical requirements. ### Geographic Focus A fund focused on a single city or region operates differently than a global fund. Hyper-local funds benefit from dense networks. They know everyone in the ecosystem. They track local events, community gatherings, and founder meetups. Deals come from repeated interactions and warm introductions within a tight network. Global funds face different challenges. Partners operate across time zones. The fund deals with multiple currencies and potentially multiple legal entities for regional investments. Different regions have different legal and compliance requirements. Regional partners might have deal attribution and carry allocation based on geography. Sourcing is also harder at scale. Global funds compete against local investors with deeper regional networks. Building relationships with founders across time zones requires more deliberate effort and often more reliance on data-driven sourcing to compensate for thinner local presence. From a technical perspective, geographic focus shapes your architecture significantly. Multi-entity structures require careful data modeling: which legal entity made which investment? How do you aggregate performance across entities while maintaining separation for legal and tax purposes? Currency handling affects portfolio valuation and financial reporting. Time zones affect scheduling features, notification timing, and how you display dates. Access control becomes more complex: regional partners might only see deals in their geography, while HQ needs global visibility. Examples of funds with a geographic focus: * [byFounders](https://www.byfounders.vc/) - Nordic focus * [Cusp Capital](https://www.cuspcapital.com/) - Continental Europe * [Blackbird](https://blackbird.vc/) - Australia and New Zealand ### Industry Specialization Funds that specialize in specific industries need tools tailored to those sectors. Deep tech and hard tech funds face long development timelines. Companies might spend years in R\&D before having a product. Diligence involves technical experts, often PhDs who can evaluate the science. The fund needs to track grants and non-dilutive funding sources. IP and patent monitoring becomes important. Success metrics differ from typical startups. Fintech and crypto funds deal with regulatory complexity. Banking licenses, compliance requirements, and regulatory approval workflows need tracking. Cap tables might involve token economics alongside traditional equity. Diligence processes need to evaluate banking partnerships and infrastructure relationships. The fund needs specialized integrations, maybe with blockchain explorers or regulatory databases. B2B SaaS funds focus heavily on financial metrics. ARR, net retention, customer acquisition cost, lifetime value. These metrics are well-defined and comparable across portfolio companies. Customer reference tracking matters. Product and technical architecture evaluation is important. Go-to-market strategy gets scrutinized. Consumer and marketplace funds care about user growth, retention, cohort analysis, and unit economics. Brand strength matters. Community and network effects need assessment. The metrics are different from B2B, and the diligence process focuses on different aspects. From a technical perspective, industry specialization shapes your sourcing and research infrastructure especially. A deep tech fund needs access to patent databases, grant tracking systems, and academic research. A fintech fund integrates with regulatory databases and compliance monitoring. A crypto fund might pull data from blockchain explorers and token analytics platforms. Your deal sourcing changes - defense tech funds track government contract databases, biotech funds monitor clinical trials, and B2B SaaS funds benchmark against industry ARR data. Examples of industry-focused funds: * [Scout Ventures](https://www.scoutventures.com/) - Defense and dual-use technology * [Paradigm](https://www.paradigm.xyz/) - Crypto and web3 * [QED Investors](https://qedinvestors.com/) - Fintech ### Generalist Funds Generalist funds invest across sectors and stages. They rely on broader pattern recognition rather than deep domain expertise in one area. Portfolio construction is deliberately diverse. Team members bring varied backgrounds. The technical challenge for generalist funds is flexibility without chaos. Your categorization systems need to handle companies that don't fit neat boxes. Tagging and metadata become more important than rigid hierarchies. Signal processing for deal flow requires broader pattern matching: you can't rely on sector-specific signals, so you need systems that surface opportunities based on founder quality, market dynamics, and thesis fit across different industries. Your metric tracking needs to gracefully handle B2B SaaS, consumer apps, marketplaces, and hardware companies in the same portfolio. The key is building abstractions that work across sectors while allowing customization where needed. Examples of generalist funds: * [Atomico](https://atomico.com/) - European multi-stage * [Creandum](https://creandum.com/) - European early stage * [Bessemer Venture Partners](https://www.bvp.com/) - US multi-stage generalist ## Understanding Your Fund's Thesis ### "Show Me Your Portfolio and I'll Show You Your Thesis" A fund's stated thesis and their actual thesis often differ. The pitch deck to LPs might say "Series A B2B SaaS" but the portfolio tells a different story. The best way to understand what a fund actually looks for is to analyze what they've invested in. Start by looking at the portfolio data. What stages are they actually investing at? If they say Series A but 30% of investments are seed rounds, that tells you something. What industries keep appearing? Is there a pattern you can identify, or is it truly diverse? What check sizes are common? Look at the range and average. What geographies show up? Are they concentrated or distributed? What patterns exist in founding teams? Do they back repeat founders, or first-timers? Then look for exceptions. Which investments don't fit the obvious pattern? Were these experiments, strategic bets, or signs of thesis evolution? Sometimes a fund makes one unusual bet that reveals a new direction. Did the strategy shift between Fund I and Fund II? Funds evolve, and understanding that evolution helps you understand where they're going. Finally, understand the trajectory. Is the fund moving up-market, writing larger checks at later stages? Are they expanding into new sectors or focusing more narrowly? These trends inform what features matter most. Here's an example. A fund says: "We're a Series A fund focused on B2B SaaS." But when you analyze the portfolio, you see 60% Series A, 30% Seed, 10% Series B. The industry breakdown is 70% B2B SaaS, 20% fintech, 10% consumer. Average check is \$3M with a range from \$500K to \$8M. What does this tell you? They're flexible on stage. You need to build for seed through Series B, not just Series A. B2B SaaS is primary but not exclusive. Your categorization and metrics need flexibility. The wide check size range suggests either follow-on investing or opportunistic deals at different stages. The thesis is more nuanced than the tagline, and your software needs to reflect that nuance. ### Meeting with GPs: The Right Questions To understand what to build, you need to understand how the GPs actually work. Interviews reveal workflows, pain points, and data needs. Here are questions that reveal technical requirements: **About Thesis Creation and Evolution:** * "How do you develop and refine your investment thesis?" * "What data or signals inform changes to your thesis over time?" * "How do you track emerging trends or technologies you're monitoring?" * "How do you share thesis work across the team?" **About Deal Flow:** * "Walk me through how a deal comes in and what happens next" * "How many deals do you see per month? How many do you invest in?" * "What's your typical timeline from first meeting to close?" * "Who needs to be involved in decision-making?" * "What opportunities have you missed that you wish you hadn't? Why did you miss them?" **About Due Diligence:** * "What does your diligence process look like?" * "What information do you need to make a decision?" * "How do you share information within the team?" * "What are the common blockers that slow down deals?" * "When you pass on a company, how do you track why you decided not to invest?" **About Portfolio Management:** * "How often do you interact with portfolio companies?" * "What information do you need from them regularly?" * "How do you decide on follow-on investments?" * "What board materials do you receive?" **About LP Reporting:** * "How often do you report to LPs?" * "What do LPs ask for that's hard to produce?" * "How long does it take to prepare LP reports?" * "What data is hardest to gather?" **About Pain Points:** * "What takes the most time that shouldn't?" * "What information do you wish you had but don't?" * "What manual processes drive you crazy?" * "What tools have you tried that didn't work? Why?" The goal isn't to gather feature requests. It's to understand workflows, pain points, and data needs before proposing solutions. ### Observation Over Interviews Interviews tell you what people think they do. Observation shows you what they actually do. There's always a gap. Shadow the team whenever possible. Sit in on partner meetings. Watch how deals get discussed. What information do they reference? What questions get asked? See how they prepare for IC. What documents get created? What analysis happens? Observe portfolio company interactions. How do updates get shared? What matters in those conversations? Watch the current tools in use. What spreadsheets do they maintain? These often reveal the data they care about most. What emails get sent repeatedly? Look for templates and patterns. What Slack channels are most active? This shows where conversation and collaboration happen. What questions get asked over and over? These reveal information gaps. Look specifically for workarounds. These are gold for understanding what's broken. Manual processes that should be automated reveal workflow inefficiencies. Information that lives in someone's head rather than a system shows where knowledge management fails. Data that gets re-entered multiple times indicates integration gaps. Decisions made without data or with incomplete data show where analytics and reporting fall short. Watch for red flags in your conversations. If someone says "We'll figure out the process as we build the tool," that's a sign you don't understand their workflow yet. If they say "Just build something flexible and we'll adapt," they don't actually know what they need. If they say "Make it like \[other fund's tool] but better," they're guessing about what might work. If they say "We need everything that \[vendor] has," they haven't thought about their actual requirements. These red flags mean you don't understand the fund well enough yet. Keep observing, keep asking questions, and don't start building until you have real clarity. ## Translating Thesis to Technical Requirements Once you understand the fund's thesis and actual operations, you can translate that into technical decisions. The seven questions above might feel abstract, so let's see how they shape real technical decisions with two contrasting examples. These examples show how fund characteristics map directly to technical architecture. Same industry (VC), completely different technical requirements. ### Example 1: Concentrated Seed Fund Imagine a fund with these characteristics: They make 5-8 investments per year. Check sizes are \$500K-\$1.5M. They review 300-400 companies annually to find those investments. The diligence process takes 2-3 weeks. They focus on deep tech and take concentrated bets on technical founders. They reserve 20% of the fund for follow-ons. This translates into specific technical requirements. The deal flow tool needs to support research-driven sourcing. Track companies the fund has researched proactively, not just inbound. Tag companies by technical area or thesis fit. Track founder conversations over time since relationships often develop over months before an investment. Support a pipeline that moves from research to initial meeting to evaluation to decision. The diligence process is quick but needs structure. Basic checklists for technical evaluation, founder reference calls, and market assessment. Document storage for pitch decks, technical assessments, and reference notes. The approval workflow should be simple but clear. Maybe partner discussion followed by consensus decision. Portfolio tracking needs to be engaged but not overwhelming. Track key metrics for early-stage companies: burn rate, runway, revenue, technical milestones, hiring progress. Support regular check-ins with founders. Track follow-on investment decisions. Which companies are raising Series A? Do we want to invest more? What's the reserved capital allocated to each company? The fund likely doesn't take formal board seats at seed but maintains close relationships. Track advisory board participation, regular founder calls, and value-add activities. Integration with portfolio company tools helps. Maybe they use a platform where companies report metrics quarterly. LP reporting happens quarterly with narrative depth. Not just numbers, but the story of each investment. What progress has the company made? What are the risks and opportunities? What's the plan for follow-on investment? What not to build: Don't build for processing hundreds of deals per month. This fund is selective, not volume-driven. Don't assume minimal engagement. Seed funds care deeply about their portfolio companies and stay closely involved. Don't skip follow-on investment tracking. Many seed funds reserve significant capital for follow-ons and need to manage that actively. ### Example 2: Series A Fund with Deep Diligence Now imagine a different fund: 20-30 deals per year. Check sizes of \$3M-\$10M. Diligence process takes 3-6 weeks involving multiple team members. The fund always takes board seats and plans to be actively involved. They're generalist, investing across sectors. This requires completely different technical architecture. The deal flow tool needs structure and collaboration. Multi-stage pipeline tracking deals from sourcing through diligence to IC to closing. Each stage has different requirements and different people involved. Deal team assignment and workload management matters. Who's working on what? Are deals distributed evenly? Being generalist makes sourcing particularly challenging. The team sees hundreds of companies across different industries. You need systems to sift through the noise and surface the right opportunities. Signal scoring based on thesis fit, founder quality, market dynamics. Recording conversations becomes critical since partners have dozens of calls with founders and advisors. Entity resolution matters too: is this the same company we saw last year, or the founder's previous startup? Diligence checklist and document management become important. What needs to be done for each deal? Where are the documents? IC presentation preparation needs support. Assembling materials from multiple sources into a coherent partner presentation. Portfolio tracking needs to be detailed and operational. Comprehensive metrics across financial, product, and team dimensions. Not just revenue, but also unit economics, customer metrics, product development, hiring progress. Board meeting scheduling and materials management matters when you have 20 board seats. Follow-on investment analysis requires tracking reserved capital and making data-driven decisions about when to invest more. Co-investor relationship tracking helps manage syndication. LP reporting happens quarterly with detail. Portfolio company deep-dives where you tell the story of what's happening, not just numbers. Automated data aggregation from multiple sources because manually assembling quarterly reports is painful. Performance attribution so LPs understand which investments are driving returns. What not to build: Don't treat every inquiry as a serious deal. You need triage. Most inbound shouldn't enter the full pipeline. Don't use one-size-fits-all metrics. A B2B SaaS company and a marketplace need different KPIs. The system should support flexible metrics per company. Don't rely on manual board meeting management. With 20+ portfolio companies, spreadsheets and email don't scale. ## The Bottom Line Before you write any code, get clear on seven critical questions: **Stage**: What stage does this fund invest at? Pre-seed, Seed, Series A/B, Growth, or PE? This determines complexity, workflows, and data depth. **Volume**: What's the deal volume? 100+ deals per year versus 10 deals per year changes everything about how you architect the system. High volume needs speed and filtering. Low volume needs depth and detail. **Strategy**: What's the investment strategy? Geo-focused, sector-specific, or generalist? This determines customization needs, specialized features, and integration requirements. **Involvement**: How hands-on is the fund? Do they take board seats? Provide operational support? Remain passive? This shapes portfolio management features and the depth of portfolio company tracking. **Team**: What does the team look like? Solo GP, small team, or large organization? Team size impacts workflow complexity, collaboration needs, and user permissions. **Thesis**: What do they actually invest in? Don't trust the pitch deck. Look at the portfolio. Analyze patterns. Understand the nuanced reality of what gets funded. **Operating Style**: What's the organization's culture? Move fast and break things, or stable release cycles? Is the tech team embedded with deal teams or a separate function? This determines your deployment approach, testing requirements, and how you prioritize features. These seven answers determine your entire technical architecture. There's no one-size-fits-all solution. The most expensive mistake you can make is building generic VC software and trying to adapt it later. Start specific. Understand your fund deeply. Build exactly what they need. Expand carefully based on actual requirements, not hypothetical futures. In the next chapter, we'll cover the common mistakes that even experienced developers make when starting at a VC fund, and more importantly, how to avoid them before they cost you months of wasted work. # What is a VC Fund? Source: https://buildingfor.vc/guide/part-1-understanding-vc/what-is-a-vc-fund Fundamental terms, structures, and concepts that apply across all VC funds - your glossary and mental model for how venture capital works. ## Overview Before building software for venture capital, you need to speak the same language as investors. This chapter covers the fundamental terms, structures, and concepts that apply across all VC funds. Think of this as your glossary and mental model for how VC works. You don't need to be an expert investor, but you need to understand the basic mechanics well enough to make good technical decisions. ## Resources * **[The Business of Venture Capital, 3rd Edition](https://www.amazon.com/dp/1119639689)** by Mahendra Ramsinghani - Comprehensive overview of fund structures, legal frameworks, and the LP-GP relationship * [**Venture Deals**](https://www.amazon.com/dp/1119594820) by Brad Feld and Jason Mendelson - Essential reading on how VC deals work * [**Secrets of Sand Hill Road**](https://www.amazon.com/dp/0753553961) by Scott Kupor - Inside look at how VC firms operate * [**Moonfare Glossary**](https://www.moonfare.com/glossary) - Comprehensive glossary of private equity and venture capital terms ## The Players: LPs, GPs, and Portfolio Companies At its core, a venture capital fund involves three groups: Limited Partners who provide the capital, General Partners who manage it, and Portfolio Companies who receive it. Understanding these relationships is fundamental to everything else. ### Limited Partners: The Money Limited Partners, or LPs, are the investors in a venture capital fund. They're the ones providing the capital that eventually gets invested in startups. LPs come in many forms: pension funds managing retirement savings for teachers and firefighters, university endowments looking to fund scholarships, family offices managing wealth for wealthy families, and sometimes corporations looking to gain strategic exposure to emerging technologies. The word "limited" is important here. LPs have limited involvement in the fund's day-to-day operations. They commit capital to the fund but don't make investment decisions or sit in partner meetings debating which startup to back. Their role is largely passive: they commit money, receive regular reports on how that money is being deployed, and eventually receive distributions when portfolio companies exit. This limited involvement comes with significant confidentiality. LP identities and their commitment amounts are typically kept private. An LP might not want it publicly known that they're investing in venture capital, or they might not want other GPs to know the size of their commitments elsewhere. This confidentiality shapes everything from how data is secured to who can access what information in any software you build. It's also crucial to understand that when an LP commits \$10 million to a fund, they don't write a \$10 million check on day one. Instead, they promise to provide that capital over time as the fund needs it for investments. The money stays in the LP's accounts earning returns until the fund calls for it. This distinction between committed capital and called capital runs through every aspect of how VC funds operate. ### General Partners: The Operators On the other side are the General Partners, or GPs. These are the investors you think of when you picture venture capitalists: the partners who source deals, conduct due diligence, make investment decisions, sit on boards, and work with portfolio companies. They're the ones running the fund's operations day to day. Unlike LPs, GPs make all the decisions about how the fund's capital gets deployed. They decide which companies to invest in, how much to invest, what terms to negotiate, and when to exit. GPs make money in two ways. First, they receive management fees, typically 2% of the committed capital per year. For a \$100 million fund, that's \$2 million annually to cover salaries, office space, travel, and all the operational costs of running the fund. These fees are paid regardless of how well the investments perform. Second, and more importantly, GPs receive carried interest, or "carry." This is their share of the profits, typically 20%. But here's the key: carry only gets paid after the LPs have received back all their invested capital, and sometimes only after a preferred return hurdle is met. This alignment of incentives means GPs only make serious money if the fund performs well for the LPs. The way GPs spend their time reveals what any software system needs to support: sourcing deals through their networks, evaluating companies, conducting due diligence, preparing for and running investment committee meetings, negotiating terms with founders, sitting on portfolio company boards, and maintaining relationships with LPs. Each of these activities generates data and requires workflows. ### Portfolio Companies: The Investments Portfolio companies are the startups and businesses that receive investment from the fund. They're the reason the whole ecosystem exists: LPs provide capital to GPs so GPs can invest in promising companies that will hopefully generate returns. From the fund's perspective, portfolio companies exist in different states. Before investment, they're prospects moving through the deal pipeline. They get evaluated, researched, debated. After investment, they become part of the portfolio. They get tracked, monitored, supported. The relationship changes from courtship to partnership. Once invested, the fund's relationship with a portfolio company is ongoing and multifaceted. Most funds that lead rounds or make significant investments take a board seat, giving them formal governance responsibility and requiring regular attendance at board meetings. Even without a board seat, funds want regular updates: quarterly financial reports, monthly metrics, informal check-ins about challenges and wins. Portfolio companies send data to their investors constantly: revenue numbers, burn rate, cash runway, customer acquisition metrics, hiring updates, product milestones. The format varies wildly. Some funds have structured data collection through portals or templates, others receive casual email updates with PDFs attached. Some companies are religious about monthly updates; others go quiet for months then surface when they need help or are raising another round. The fund also provides value beyond capital. GPs make introductions to potential customers, help recruit executives, advise on strategy, connect companies to later-stage investors, and sometimes help navigate crises. This value-add varies dramatically by fund; some are highly engaged, others are passive financial investors. But even passive funds need to track their portfolio companies' progress to report to LPs and make follow-on investment decisions. Portfolio companies also create most of the complexity in a fund's data model. Each company goes through multiple financing rounds at different valuations. Ownership percentages change over time as new investors join and existing investors get diluted. Companies pivot, changing their business model and metrics. Some succeed spectacularly, some fail, most are in between. Some exit through acquisitions, some through IPOs, some through secondary sales, some die quietly. Understanding portfolio companies means understanding that they're not static database records. They're dynamic businesses that evolve over years. The company you invested in at seed stage looks completely different three years later when they've raised a Series B, hired 50 people, and pivoted twice. Your system needs to track this evolution while maintaining the historical record of how the company has developed. ## The Structure: Companies and Vehicles Here's where things get confusing for most people new to VC: a "venture capital fund" isn't just one entity. It's actually a collection of related but legally separate entities, each serving a specific purpose. ### The Management Company The management company is the operating entity. Think of it as the business itself. If you see "Acme Ventures LLC," that's probably the management company. This is the entity that employs the team, signs the office lease, and pays the bills. The management company receives management fees from the fund and uses those fees to cover operating expenses. A single management company typically manages multiple funds over time. Acme Ventures LLC might manage Acme Ventures Fund I, then raise and manage Acme Ventures Fund II a few years later, and so on. The same team, the same entity, but different fund vehicles. ### The Fund Vehicle The actual fund (let's say "Acme Ventures Fund I, LP") is a separate legal entity, usually structured as a limited partnership. This is the vehicle that holds the LP commitments, makes the investments, owns the equity in portfolio companies, and distributes returns. The LPs are limited partners in this entity, and the management company (or its affiliates) serves as the general partner. This separation isn't just legal technicality. Money flows differently through these entities. Management fees flow from the fund to the management company. Investment capital flows from LPs to the fund to portfolio companies. Returns flow back from portfolio companies to the fund to LPs. Understanding these flows is essential because you'll need to track them separately, allocate expenses correctly, and report on them accurately. ### SPVs and Side Cars As if two entities weren't enough, there are often more. Special Purpose Vehicles, or SPVs, are created for specific investments. Let's say Acme Ventures wants to invest \$5 million in a hot company, but several LPs want additional exposure. The fund might create an SPV just for this deal, allowing those LPs to invest beyond their fund commitment. It's a separate legal entity, created for one investment, with its own investors and terms. Side cars are similar but broader. They're parallel vehicles that invest alongside the main fund across multiple deals. A fund might create a side car for a specific LP who wants more exposure than their fund commitment allows, or for a group of co-investors who want to invest in the fund's deals. Each of these entities needs to be tracked separately, but they're all interconnected. An investment might involve capital from the main fund, an SPV, and a side car. Ownership calculations need to account for all three. Reporting needs to aggregate across them appropriately. It's complex, and it's one of the first places developers go wrong if they don't understand the structure. ## How Money Moves: The Capital Lifecycle The movement of capital through a VC fund follows a predictable but often misunderstood lifecycle. It's not as simple as "LPs give money, fund invests money, fund returns money." The reality is more nuanced and plays out over years. ### Fundraising and Commitments It starts with fundraising. The GPs go out and pitch their strategy to potential LPs: "We're raising a \$100 million fund to invest in early-stage B2B SaaS companies in Europe." LPs who are interested don't write checks immediately. Instead, they sign commitment letters promising to provide capital when called upon. A fund might have a target size of \$100 million but a hard cap of \$150 million. They'll announce a "first close" when they reach the minimum needed to start investing, say \$50 million in commitments. They can now legally start making investments. The fund continues fundraising until they reach final close, which might be \$120 million in total commitments. This fundraising period typically takes 6-18 months. During this time, LPs haven't actually transferred most of their money yet. They've just promised it. This is a crucial distinction that confuses many people building VC software: committed capital is not the same as cash in the bank. ### Capital Calls When the fund finds an investment opportunity (let's say they want to invest \$3 million in a promising startup), they don't have a bank account with \$100 million sitting in it. Instead, they issue a capital call to the LPs. A capital call is a formal notice saying "We need \$3 million for this investment. Each of you needs to send us your pro-rata share within 30 days." An LP who committed \$10 million (10% of the fund) would need to wire \$300,000 (10% of the call). The fund collects the capital from all the LPs, then wires the money to the portfolio company and receives equity in return. This happens over and over throughout the investment period of the fund. Each time the fund makes an investment, it calls capital. This means the fund is constantly tracking how much each LP has committed, how much has been called, and how much remains uncalled (often called "dry powder"). LPs occasionally miss or delay capital calls, which creates its own complications. The fund needs to track who's paid, who hasn't, follow up with delinquent LPs, and handle the paperwork. It's a workflow that software needs to support reliably. ### Making Investments When the fund receives the called capital, it invests in a portfolio company. But "investment" isn't as simple as it sounds. The fund might receive common equity, preferred equity, a SAFE (Simple Agreement for Future Equity), a convertible note, or various other instruments. Each has different terms, conversion mechanics, and implications for ownership. The fund tracks its cost basis (the actual dollars invested) but also needs to track current ownership percentage, which can change over time as the company raises more money and the fund gets diluted. The fund might invest in the company multiple times across different rounds, each at different valuations. Tracking all of this accurately is essential for calculating fund performance and reporting to LPs. ### Portfolio Management After investing, the fund monitors the portfolio company. Some GPs sit on the board, attending meetings quarterly or monthly. Portfolio companies send updates. Sometimes these are structured financial reports, sometimes casual email updates. The fund tracks key metrics: revenue, burn rate, runway, user growth, whatever matters for that particular company and sector. The fund also decides whether to invest more. Most VC funds reserve capital for follow-on investments in their best companies. When a portfolio company raises a Series B, the fund needs to decide: do we invest again to maintain our ownership percentage (exercising pro-rata rights), or do we let ourselves get diluted? This decision-making process involves data, discussion, and usually formal approval. ### Exits and Distributions Eventually, some portfolio companies exit. A company gets acquired, goes public, or sells shares in a secondary transaction. The fund receives either cash or liquid stock in return for its equity. This is when LPs finally see returns. But the money doesn't just get distributed immediately. First, it flows through what's called a distribution waterfall. The standard VC structure works like this: first, return all the capital that LPs invested. If the fund called \$80 million total from LPs over the fund's life, the first \$80 million of exits goes back to LPs. Then split remaining profits: typically 80% to LPs, 20% to GPs as carry. Some funds (more common in PE) include a preferred return hurdle between these steps, typically 8% annually that LPs must receive before carry is paid. These waterfall calculations get complex fast. Different investments exit at different times. LPs care about exactly when they receive distributions. The fund needs to track everything precisely: how much each LP contributed across all capital calls, how much they've received in distributions so far, and how profits should be split according to the fund's specific waterfall structure. Distributions happen throughout the fund's life as companies exit, not all at once at the end. A fund might make its first distribution in year 3, continue distributing through year 10, and potentially make final distributions in year 12 or 15 if some companies took a long time to exit. Throughout all of this, the fund reports to LPs on performance using metrics like DPI, TVPI, and IRR (which we'll define shortly). ## The Language: Key Terms Every field has its jargon, and venture capital is no exception. Here are the terms you'll hear constantly and need to understand deeply if you're building software in this space. ### Performance Metrics **[DPI (Distributions to Paid-In Capital)](https://www.moonfare.com/glossary/distributed-to-paid-in-capital-dpi)** is the most important metric for LPs because it measures actual cash returned. Take the total cash distributed to LPs and divide by the total capital called from LPs. If a fund called \$80 million and has distributed \$120 million, the DPI is 1.5x. This is real money in LPs' bank accounts, not theoretical value. **[TVPI (Total Value to Paid-In Capital)](https://www.moonfare.com/glossary/total-value-to-paid-in-capital-tvpi)** includes both distributions and the current value of remaining investments. If that same fund has distributed \$120 million and its remaining portfolio companies are valued at \$40 million, the TVPI is 2.0x (\$160M / \$80M). This metric is less reliable than DPI because it depends on valuations of unrealized investments, which might not materialize. **[IRR (Internal Rate of Return)](https://www.moonfare.com/glossary/internal-rate-of-return-irr)** is the annualized return accounting for the timing of cash flows. It's time-weighted, so returning 3x in 3 years is much better than returning 3x in 10 years, and IRR captures that difference. The calculation is complex. It's finding the discount rate where the net present value of all cash flows equals zero. **[MOIC (Multiple on Invested Capital)](https://www.moonfare.com/glossary/multiple-on-invested-capital-moic)** is the simplest: total value divided by total invested. MOIC and TVPI both measure investment performance, but differ in their denominators: MOIC divides total value by invested capital, while TVPI divides total value by paid-in capital (including fees). ### Investment Instruments Funds don't always receive simple equity. Different instruments have different terms, conversion mechanics, and implications. **Equity (Preferred Stock)** is direct ownership in the company. When a fund invests in a Series A, they typically receive preferred stock with specific rights: liquidation preferences, board seats, protective provisions, anti-dilution protection. This is straightforward ownership but comes with negotiated terms that affect outcomes in exits. **Common Stock** is what founders and employees typically own. It has no special preferences or protections. Investors rarely take common stock except in unusual circumstances. **SAFE (Simple Agreement for Future Equity)** is a convertible instrument popular in early-stage investing, invented by [Y Combinator](https://www.ycombinator.com/documents). The investor gives the company money now, and the SAFE converts into equity later during a priced round. SAFEs typically have a valuation cap (the maximum valuation at which they convert) and/or a discount (investors get a better price than the next round). They're "simple" because they defer valuation negotiations until later. **Convertible Note** is similar to a SAFE but structured as debt. The investor loans money to the company with interest, and the note converts to equity in a future priced round. Like SAFEs, they typically have a cap and/or discount. The key difference is that convertible notes have a maturity date and accrue interest, though in practice they almost always convert rather than being repaid. Why this matters: Each instrument type needs different data modeling. A SAFE doesn't give you ownership percentage until it converts. A convertible note has interest calculations. Tracking what will happen when instruments convert requires understanding caps, discounts, and the terms of future rounds. ### Funding Rounds Venture capital happens in stages, each with different characteristics, valuations, and investor types. **Pre-Seed** is the earliest stage, often before the company has launched or has meaningful revenue. Rounds are typically \$100K-\$1M, often from angel investors, very early stage funds, or accelerators. Companies might just have a prototype or idea. High risk, very early. **Seed** is when the company has proven some concept but is still early. Rounds are typically \$1M-\$5M. The company might have initial product-market fit, some revenue, or strong early traction. Seed funds specialize in this stage, along with angel investors and some multi-stage funds doing early deals. **Series A** is the first institutional round for most companies. Rounds typically \$10M-\$20M. The company has proven product-market fit and needs capital to scale. They usually have meaningful revenue or user growth. Series A funds look for clear metrics: revenue growth, customer acquisition, team strength. This is where things get more formal: boards, proper governance, structured metrics. **Series B** is about scaling what works. Rounds typically \$15M-\$50M. The company has proven they can grow and now needs capital to expand: new markets, more sales people, larger team. Clear business model, strong revenue, maybe even profitability or a clear path to it. **Series C and beyond** are later stages focused on scaling faster or expanding into new areas. Rounds can be \$50M-\$200M+. These companies are usually mature businesses with strong revenue, clear unit economics, and paths to exit. The focus shifts from "will this work?" to "how big can this get?" **Growth/Late Stage** rounds are the final stages before IPO or acquisition. Rounds can be \$100M-\$500M+. Companies are often profitable or nearly so, with hundreds of millions in revenue. These look more like PE deals than traditional venture. Why this matters: The stage determines everything about how you track the investment. Seed investments might not have board seats or detailed metrics. Series B+ investments require comprehensive tracking, board management, and detailed reporting. Your software needs to handle the complexity appropriate to each stage. Note on round sizes: These ranges are illustrative and vary significantly by geography, sector, and market conditions. US rounds tend to be larger than European rounds, and AI-focused companies have pushed averages higher since 2023. For current benchmarks, see the [PitchBook-NVCA Venture Monitor](https://nvca.org/wp-content/uploads/2026/01/q4-2025-pitchbook-nvca-venture-monitor.pdf) or [Crunchbase](https://news.crunchbase.com/venture/funding-rounds-average-mean-startups-charts/) funding data. ### Valuation Terms When a fund invests in a company, the **pre-money valuation** is what the company was worth before the investment. If a company has a \$10 million pre-money valuation and the fund invests \$3 million, the **post-money valuation** is \$13 million. The fund owns \$3M / \$13M = 23% of the company. This math is fundamental to tracking ownership. **Pro-rata rights** give investors the option to invest in future rounds to maintain their ownership percentage. If the fund owns 23% and the company raises a Series B, pro-rata rights let the fund invest enough to stay at 23% instead of getting diluted down. These rights are negotiated as part of each investment and need to be tracked because they affect follow-on investment decisions. **Liquidation preferences** determine who gets paid first in an exit. A "1x liquidation preference" means the fund gets their money back before other shareholders see anything. In a \$50 million exit, if the fund invested \$10 million with a 1x preference, they get \$10 million off the top, and the remaining \$40 million is split among all shareholders. More complex preferences (2x, participating preferences) exist and dramatically affect exit proceeds. **Dilution** is what happens when a company issues new shares, reducing everyone's percentage ownership. If the fund owns 20% and the company raises a new round, the fund might get diluted down to 15%. Tracking dilution accurately requires understanding every financing event in the company's history and how it affected the cap table. ## How GPs Actually Make Money It's worth understanding the economics of running a VC fund, because they shape everything about how funds operate and what they prioritize. ### Management Fees: Keeping the Lights On The standard structure is 2% of committed capital annually, though this varies. For a \$100 million fund, that's \$2 million per year. Over a typical 10-year fund life, that's \$20 million total. Many funds step down fees after the investment period ends - 2% for the first 5 years while actively investing, then 1.5% or less while just managing existing investments. This fee revenue needs to cover everything: partner salaries, analyst salaries, office rent, travel, legal fees, accounting, software, data, subscriptions, events, everything. For a small fund, 2% often isn't enough to cover a large team, which is why many seed funds are small teams or solo GPs. For larger funds, management fees can support bigger teams with specialized roles. ### Carry: The Real Upside Carried interest is where GPs make real money, but only if the fund performs. The standard is 20% of profits, but this only kicks in after returning LP capital and paying preferred return (if applicable). Here's an example: a \$100 million fund fully deploys its capital into portfolio companies. Years later, the portfolio exits for \$300 million total. First, \$100 million goes back to LPs (their capital back). That leaves \$200 million in profit. LPs get 80% of the profit (\$160 million), and GPs get 20% (\$40 million) as carry. That \$40 million carry is split among the GPs according to the partnership agreement. A large fund with many partners might split it widely; a small fund with two GPs might split it 50/50. Either way, carry is the incentive to generate returns, and it dwarfs management fees for successful funds. ## The Process: How Deals Flow Understanding how a deal moves through a fund from first contact to portfolio company helps explain what data needs to be tracked and what workflows need to be supported. It starts with **research and thesis development**. Many funds don't just wait for deals to come to them. They proactively research macro trends, emerging technologies, and market dynamics to build investment theses. A fund might spend months researching AI infrastructure, talking to experts, mapping the ecosystem, and developing a point of view on where opportunities exist. This research informs what they look for and where they focus their sourcing efforts. Then comes **sourcing**. Deals come in through warm introductions from other founders, VCs, or friends in the network. They come from cold emails to a fund's info@ address. They come from events, demo days, conferences. But increasingly, deals also come from proactive outreach based on research. A fund identifies interesting companies from their market mapping and reaches out directly. Someone at the fund (often a junior team member) tracks all these incoming leads and flags which came from proactive research versus inbound. Then comes **screening**, the fast filter. Does this even fit the fund's thesis? Right stage, right geography, right sector? A quick look at the deck, maybe a 15-minute call. Most deals get rejected at this stage with a polite pass. It's about volume management. A fund might see 500 deals a year and screen out 450 of them quickly. For deals that pass screening, **due diligence** begins. This is where the fund does real work: market research, customer reference calls, technical evaluation if it's a technical product, financial model review, background checks on founders, legal review of the company structure and contracts. Different funds have different diligence processes, from lightweight for seed investments to extensive for growth equity deals lasting months. The **investment committee** (IC) is where investment decisions are made. Someone at the fund (the deal champion) presents the opportunity to the partnership. There's discussion, debate, questions. The committee decides: pass, invest, or "not now but circle back later." The decision might be subject to conditions: "yes, but only if we can lead the round" or "yes, but at this valuation, not their asking price." **Syndication** often happens next, especially at earlier stages. The lead investor rarely fills the entire round alone. They bring in other funds, angels, or existing investors to complete the round. This involves coordinating with other parties, sharing deal information (with founder permission), and sometimes negotiating allocation if the round is oversubscribed. At pre-seed and seed, rounds are frequently syndicated across many investors; at later stages, rounds typically have fewer but larger investors. The fund needs to track who else is in the deal and coordinate timing for closing. If the IC says yes, the **closing** process begins. Lawyers draft and negotiate documents. Terms get finalized. The company, fund, and other investors sign the paperwork. The fund issues a capital call to LPs. Money is wired. Equity is issued. The company is now officially in the portfolio. Finally, **post-investment** begins. The fund tracks the portfolio company's progress, supports them where possible, attends board meetings if they have a seat, participates in future financing rounds, and eventually helps guide the company toward an exit. This phase lasts years and generates most of the ongoing work for a fund. Each stage has different data needs, different participants, and different workflows. A deal moves through these stages over weeks or months, and the fund needs to track where each deal is, who's responsible, what's been done, and what needs to happen next. ## The Timeline: Fund Lifecycle VC funds operate on long timelines, typically 10 years with options to extend. Understanding this lifecycle helps explain why different features matter at different times. **Year 0** is fundraising. The GPs are pitching LPs, negotiating terms, signing commitment letters. No investing is happening yet because there's no fund yet. Once they reach first close (the minimum threshold), the fund legally exists and investing can begin. **Years 1-4** are the investment period. The fund is actively looking at deals, making new investments, deploying capital. This is when deal flow management matters most. The fund is building its portfolio, typically making 20-50 investments depending on stage and strategy. Capital calls go out regularly as new investments are made. **Years 5-10** are the harvesting period. The fund has largely stopped making new investments except follow-ons in existing portfolio companies. The focus shifts to helping portfolio companies grow and exit. As companies exit, distributions flow to LPs. The fund's performance metrics start becoming real as DPI increases. Reporting to LPs becomes more focused on realized returns rather than unrealized valuations. **Years 10+** are extension periods, if needed. Some portfolio companies haven't exited yet. The fund extends its life to continue managing those remaining investments until they can exit. It's not ideal (LPs prefer getting their money back sooner), but it's common in venture capital where exits take time. This timeline shapes what software features matter when. A young fund cares about deal flow and making investments efficiently. A mature fund cares about portfolio monitoring and LP reporting. A fund in harvesting mode cares about tracking exits and calculating distributions. The priorities shift as the fund ages. ## Common Misconceptions Before we tie this all together, let's clear up some common misunderstandings that lead to bad technical decisions. **"VCs have a pile of cash to invest"** - No, they have LP commitments. The cash is still in LP bank accounts until it's called. This means funds care deeply about capital call management and timing. Software that assumes cash is always available will break workflows. **"All VC funds work the same way"** - Not at all. Stage, size, and strategy create vastly different operations. A seed fund operates nothing like a growth equity fund. Software built for one won't work for the other without major changes. This is why [understanding your specific fund](/guide/part-1-understanding-vc/understanding-your-vc-fund) matters so much. **"Exits happen quickly"** - The average time from investment to exit is 7-10 years. Sometimes longer. This long timeline means data needs to be preserved, systems need to be maintained for years, and investors care about tracking companies over long periods. Features that assume quick exits will feel broken. **"VCs just pick winners"** - Portfolio construction matters more than most people realize. A fund needs enough shots on goal to find winners, but not so many that they're spread too thin. Follow-on investment strategy matters. How active the fund is post-investment matters. It's not just picking; it's managing a portfolio over time. **"Fund performance is easy to calculate"** - IRR calculations are notoriously tricky to get right. Waterfall calculations have edge cases. Unrealized valuations are subjective. Performance reporting to LPs needs to be precise because it's often audited. This is not something to build carelessly. ## Why This All Matters for Building Software Now that you understand how VC funds actually work, the technical implications become clearer. Every concept we covered translates directly into how you build software: **LPs and GPs** aren't just "users" in your system. They're different entities with different data access needs. LP data is confidential - LPs often can't see other LPs' information. GPs need collaborative access to deals and portfolio data. You need role-based access control that reflects these real-world boundaries. **The distinction between management company and fund** means you need multi-entity data models from day one. Expenses, fees, and investments all flow through different legal entities. Reports need to aggregate data correctly across entities. Getting this wrong early means painful refactoring later. **Capital calls** are not just transactions; they're workflows. You need to track who's been notified, who's paid, who's late. You need to generate official notices, track receipt of funds, handle exceptions. This is a core workflow that needs to be reliable, auditable, and compliant. **Investments over time** means you can't just track "current ownership." You need to track every round, every valuation change, every dilution event. Your data model needs to support temporal data, versioning, and the ability to recreate historical views. Questions like "what was our ownership percentage at Series A?" need to be answerable years later. **Performance metrics** like IRR and DPI need to be calculated correctly. These aren't approximate. They're used in official LP reports and audits. Getting IRR wrong by implementing a buggy formula will erode trust fast. These calculations need to be precise, testable, and documented. **The investment process** from sourcing to post-investment isn't a linear pipeline. Deals can move backward, skip stages, or sit dormant for months. Your deal flow system needs to support the messy reality of how deals actually flow, not an idealized pipeline. **The fund lifecycle** means the features you need change over time. A young fund needs robust deal flow management. An older fund needs sophisticated portfolio reporting. Your system needs to grow with the fund, or at minimum, you need to understand which features matter for your fund's current stage. **Confidentiality and compliance** aren't afterthoughts. VC data is highly sensitive. Funds have legal obligations to LPs, portfolio companies, and regulators. Audit trails, data security, access controls, and privacy protections need to be baked in from the start, not added later. Understanding these fundamentals doesn't just help you build a VC system. It helps you build the *right* VC system. You'll make better decisions about data models, you'll prioritize the right features, and you'll avoid costly mistakes that come from not understanding the domain. In the next chapter, we'll move from general VC knowledge to understanding your specific fund - because not all VC funds are created equal, and the type of fund you're building for dramatically changes what you need to build. # CRM and Deal Flow Management Source: https://buildingfor.vc/guide/part-2-tech-stack/crm-and-deal-flow Understanding why every fund needs a CRM, how to choose between options, and making it actually work for your team. ## Overview Unlike sourcing tools, which work for some funds and not others, CRM is universally valuable. Every fund tracks relationships and manages deal flow. The question isn't whether you need it, but which tool you use and whether you actually use it consistently. CRM stands for Customer Relationship Management, which is a terrible name for what venture funds need. You're not managing customers. You're tracking relationships with founders, managing deal pipeline, coordinating your team, and maintaining institutional memory about every company you've ever talked to. But "relationship and deal flow management system" doesn't fit on marketing materials, so CRM it is. The reality is that all CRMs feel terrible to use. They're not built for venture capital out of the box. They're built for sales teams with standardized processes, clear conversion funnels, and lots of repetitive activity. VC is messier. Every deal is different. The relationship with a founder you passed on three years ago might lead to your best investment next year. You can't force this into a linear sales pipeline. But building your own CRM from scratch is worse. Much worse. It's . Don't do it. Buy something customizable, integrate it with your other tools, and commit to actually using it. This chapter covers how to choose a CRM and make it work for your fund. ## Why CRMs Feel Terrible Every VC complains about their CRM. The interface is clunky. Data entry feels like homework. The workflow doesn't match how you actually work. You're constantly clicking through screens that don't apply to your process. It's tempting to think "we should just build our own." The problem is that CRMs aren't built for venture capital specifically. They're built for sales organizations. Enterprise sales has standardized stages: lead, qualified, demo, proposal, negotiation, closed. You can forecast revenue based on conversion rates. Pipeline management means moving deals through predictable stages. Venture doesn't work this way. You meet a founder. Maybe you pass. Three months later they email you an update. Six months after that you introduce them to a potential customer. A year later they're raising their next round and you lead it. Or you meet a founder, love them immediately, and invest in two weeks. The stages aren't linear. The timeline isn't predictable. The "deal" might not be a deal for years. So CRMs force you to contort your process into their model. You create custom fields. You ignore half the features. You build workarounds. It never feels quite right. This is the reality of using CRMs in venture capital. They're better than the alternative (spreadsheets, scattered notes, institutional knowledge in people's heads), but they're never perfect. The key is finding something customizable enough to match your process reasonably well, then committing to using it despite the friction. ## What You Actually Need Before choosing a CRM, understand what you actually need it to do for a venture fund. **Relationship tracking**: Every person you've met. Founders, operators, other investors, domain experts, potential LPs. How you met them. When you last talked. What you talked about. The context of your relationship. Six months from now, you should be able to look up someone and remember the full history. **Company tracking**: Every company you've evaluated. How you met them. Who on your team has relationships with the founders. The full history of your interactions over time. **Deal pipeline**: What deals (specific funding rounds) are active right now. Who's working on them. What stage they're at in your process. When decisions need to be made. What the next steps are. **Note on data modeling**: VCs invest in deals (specific funding rounds), not companies. Model your CRM at the deal level: "Acme Corp Seed Round" (Invested), "Acme Corp Series A" (Invested), "Acme Corp Series B" (Passed). If you model at the company level, you'll have problems when follow-on rounds come up - you can't mark a company as both "Invested" and "Passed" at the same time. I have first hand experience of trying to change this data model, it's a nightmare. Most good CRMs support this if you set it up correctly from the start. See [Data Modeling](/guide/part-3-technical-foundations/data-modeling) for more information. **Team coordination**: In a fund with multiple people, you need to know who's talking to which founders, who's working on which deals, and what everyone learned. The CRM prevents duplicate outreach, surfaces relevant context when someone else talked to a founder months ago, and helps the team build on each other's work. **Notes and history**: When you talked to a founder six months ago, what did you learn? When they email you an update, can you quickly see the full context? The CRM should be the source of truth for all interactions with companies and founders. **Search and filtering**: Find all the AI infrastructure companies you talked to in the last year. Find everyone you met at that conference. Find all the founders you passed on who you said you'd check back with in six months. If you can't find information quickly, you won't use the system. Unfortunately, search systems on CRMs are inexplicably poor in . **Beyond deal flow**: While this chapter focuses on deal flow management, many funds also use their CRM to manage LP relationships. Tracking interactions with current and potential LPs, managing fundraising pipelines, and maintaining investor reporting history. The same principles apply: relationship tracking, interaction history, and systematic follow-up. Some funds use separate systems for LPs, others use the same CRM with different pipelines. ## Choosing Your CRM There's no perfect CRM for venture capital. There are three main options that funds use, each with tradeoffs. **Attio**: Modern, flexible, and designed with relationship-centric businesses in mind. Better UX than traditional CRMs. Strong customization without requiring a full-time admin. Good email and calendar integration. Pricing is reasonable for small to mid-sized funds. The main limitation is that it's less established than Affinity or Salesforce, so fewer integrations and smaller ecosystem, however the APIs are easy enough to work with if you need to add some "glue code" to integrate systems. Many newer funds choose Attio because it's easier to customize to match your specific workflow without fighting the tool, it also tends to be the cheapest. **Affinity**: Built specifically for relationship-driven industries like venture capital. Automatically enriches contact data and tracks relationship strength based on email and meeting patterns. Strong at surfacing "who knows who" for warm intros. More opinionated about workflow, which means less customization but faster setup. Large funds and established GPs often choose Affinity because it has purpose-built VC workflows out of the box. The tradeoff is less flexibility. **Salesforce**: The enterprise standard. Extremely customizable. Massive ecosystem of integrations and consultants. Can be configured to do almost anything. The downsides are cost (expensive for small funds), complexity (requires real admin overhead), and UX that feels dated. Larger, established funds with dedicated operations teams sometimes use Salesforce because they need complex workflows and have resources to maintain it. Not recommended for small funds. **The decision**: Prioritize customizability over features. You'll need to adapt the CRM to match your fund's specific process. Attio and Affinity are both good choices for most funds. Attio if you want more control over structure and workflow. Affinity if you want VC-specific features. Salesforce only if you're large enough to have dedicated CRM admin resources. **Don't start with Airtable or Notion** Many funds start with [Airtable](https://www.airtable.com/) or [Notion](https://www.notion.com/) for deal flow tracking because they're familiar, flexible, and easy to set up. This is a mistake. The problem is that Airtable and Notion don't integrate well with email and calendar. You have to manually enter everything. What happens is that people stop updating it within weeks. You end up with a partially-filled database that nobody trusts or uses, and you've lost all the email and meeting history that a real CRM would have captured automatically. Just buy a real CRM from the start. The cost is worth it for the automatic email and meeting tracking alone. ## Making It Actually Work Buying a CRM isn't enough. Making it work requires integration, discipline, and actually using it. **Integrate with your other tools**: The CRM should connect to your email (Gmail, Outlook), calendar, and other tools you use. When you email a founder, it logs automatically. When you have a meeting, it logs automatically. When you add a company to your sourcing tool, it flows into the CRM. The less manual data entry, the more people will actually use it. Most CRMs have APIs. Use them. Build light integrations between your CRM and your research platform, your sourcing tools, and anything else you use. The CRM becomes more valuable when it's connected to the rest of your stack. **Use it in deal flow calls**: The single most important habit. When your team discusses deal flow, have the CRM open. Update company stages. Add notes. Assign next steps. If you talk about a company and don't update the CRM, information gets lost. Make it part of the meeting process, not something you do afterward. This also keeps the CRM current. If the system reflects the actual state of your pipeline, people trust it and use it. If it's always out of date, people stop checking it. **Commit to entering data**: There's no way around this. Someone needs to enter information about companies, founders, and meetings. With good email and calendar integration, a lot happens automatically. But initial company setup, notes from meetings that aren't on calendar, and context about relationships still require manual entry. The key is making it low-friction. If entering data takes too many clicks or too many required fields, people won't do it. Customize the system to minimize friction while capturing what actually matters. Use integrations from meeting note taking apps like [Granola](https://www.granola.ai/) to also remove friction. **Don't overthink the workflow**: Many funds spend weeks designing the perfect pipeline stages and custom fields. This is usually overengineering. Start simple. You can always add complexity later. The basics (who are we talking to, what stage are they at, what's the next step) cover 80% of what you need. The key difference from transactional sales workflows: VC deal flow isn't linear. Companies move forward, sideways, and sometimes back. You might pass on a company's seed round and lead their Series A. A "watchlist" company might sit dormant for two years then become your best investment. ```mermaid theme={null} stateDiagram-v2 direction LR [*] --> New New --> InitialContact: First meeting InitialContact --> Research: Interested InitialContact --> Watchlist: Not now InitialContact --> Passed: Not a fit Research --> TermSheet: Want to invest Research --> Watchlist: Timing wrong Research --> Passed: Decided against TermSheet --> Invested: Deal closed TermSheet --> Lost: Didn't get allocation Watchlist --> InitialContact: Revisit later InitialContact: Initial Contact TermSheet: Term Sheet ``` ## Real Example: Attio at Inflection **Author Note: Choosing a CRM** Inflection chose Attio after evaluating several options. The deciding factors were customizability and ease of integration. The setup took about a week to get the basic structure right. We kept the stages in the Deal Funnel simple: 1. **New** - Something came into our pipeline 2. **Initial Contact** - Conversed over email and had the first call 3. **Research** - Diving deep, more meetings, talking with advisors 4. **Term Sheet** - Interested and extended a term sheet 5. **Invested** - Wired payment 6. **Watchlist** - Passed for now, but could be interested again later 7. **Passed** - Actually passed 8. **Lost** - Deals we wanted but didn't get allocation This framework has worked well. The key distinctions are Watchlist (companies we might revisit) versus Passed (definitely not), and Lost (we wanted to but couldn't) versus Passed (we chose not to). These capture the non-linear nature of VC relationships better than a pure sales funnel. The hardest part wasn't the technical setup. It was the discipline of actually using it. Put it up on the screen when having the deal flow meeting. Just do it. ## The Bottom Line Buy a CRM. Don't build one. Don't start with Airtable or Notion. Choose something designed for relationship management (Attio or Affinity for most funds), customize it to match your process, integrate it with your other tools, and commit to using it. The hardest part isn't choosing the right tool. It's the discipline of actually using it consistently. Use it in deal flow calls. Make sure email and meeting tracking work automatically. Keep the data current. A mediocre CRM used consistently is better than a perfect CRM that nobody updates. In the next chapter, we'll look at fund operations software, which handles the financial side of running a venture fund. # Fund Operations Source: https://buildingfor.vc/guide/part-2-tech-stack/fund-operations Understanding fund operations software - capital management, LP reporting, and the financial infrastructure every fund needs. ## Overview Fund operations software handles the financial and administrative side of running a venture fund. Capital calls, distributions, LP reporting, portfolio valuations, compliance, and fund performance tracking. This is the infrastructure that keeps your fund running legally and financially. Many smaller funds, especially first-time fund managers, start with Excel spreadsheets. This can work for Fund I when you have a small number of LPs and a handful of investments. But as you scale into Fund II and Fund III, the migration debt piles up. You're managing multiple funds with different vintages, overlapping LPs with different commitment amounts across funds, and growing portfolio complexity. What was manageable in spreadsheets becomes error-prone and time-consuming. The right time to implement proper fund operations software is before you raise your second fund. Migrating one fund's historical data is manageable. Migrating two or three funds with years of capital calls, distributions, and portfolio changes is much harder. Unlike research platforms (which some funds might build) or sourcing tools (which some funds don't need), fund operations software is something every fund needs and should buy off the shelf. This is financial infrastructure. Don't build it yourself. This chapter covers what fund operations software actually does, how to choose between the main platforms, and how to make the transition smooth for your fund ops team. ## What You Actually Need Fund operations software handles several core functions that every venture fund needs. **Capital management**: Track capital commitments from LPs, calculate how much to call from each LP when you need to draw down funds for investments, and maintain records of who has contributed what and when. When you invest in a company, the system helps you calculate the capital call amounts and generate notices. The actual issuance of capital calls and wire transfers typically happens through your fund administrator or banking systems, but the platform tracks everything for your records. **Distributions**: When portfolio companies exit or return capital, you need to distribute proceeds to LPs according to the fund terms (waterfall calculations, carried interest, management fees). The system calculates who gets what based on your fund's specific structure and provides the distribution details. Like capital calls, the actual payments are typically handled by your fund administrator, but the platform does the math and maintains the records. **Portfolio tracking**: Track all your investments. How much you invested in each company, at what valuation, how much ownership you have, what the current valuation is, and what your position is worth. As companies raise follow-on rounds or have markups/markdowns, the system tracks the changes to your portfolio value. **LP reporting**: Generate quarterly reports for your LPs showing fund performance, portfolio company updates, capital calls and distributions, and overall fund metrics. Most LPs expect standardized reporting. The system produces these reports based on your portfolio and fund data. **Valuations and fair value**: Mark your portfolio to fair value each quarter. Early-stage investments often carry at cost until there's a markup event (new funding round at higher valuation). Later-stage funds might need more sophisticated valuation models. The system helps you track valuations consistently and document your methodology for auditors. **Fund forecasting**: Model out future fund performance. What happens if you deploy capital faster or slower? When will you need to make capital calls? When are distributions likely? What's the projected IRR and TVPI at different exit scenarios? This helps you manage the fund proactively and communicate with LPs about future expectations. **Compliance and audit**: Maintain the documentation and records that auditors and regulators require. Track management fee calculations, expense allocations, and financial statements. When your annual audit happens, the system should have everything the auditors need. **Fundraising support**: Beyond ongoing operations, these systems can generate materials useful for fundraising your next fund. Track record tear sheets, portfolio summaries, performance metrics presented in LP-friendly formats. This becomes valuable when you're raising your next fund and need to show your track record. ## Choosing a Platform Several platforms serve the venture fund operations space. All handle core functionality: the differences are in user experience, integration capabilities, and specific features. **[Carta Fund Forecasting](https://carta.com/funds/) (formerly Tactyc)**: Carta acquired Tactyc and integrated it into their platform. The advantage is tight integration if you already use Carta for fund administration or if your portfolio companies use Carta for their cap tables. Strong fund forecasting and modeling capabilities. The interface is modern and relatively intuitive. Good choice if you're already in the Carta ecosystem. **[Fundra](https://fundra.app/)**: Newer entrant focused on firm-wide adoption and flexibility. Strong on aggregating portfolio company updates automatically and making data accessible beyond the finance team. Worth evaluating if you want more customization than traditional platforms offer. **[Standard Metrics](https://www.standardmetrics.io/)**: Built specifically for institutional LPs and fund managers. Strong focus on standardized reporting and data aggregation across multiple funds. **[Vestberry](https://www.vestberry.com/)**: European-based platform popular with European funds but used globally. Strong portfolio analytics and visualization. Good LP reporting capabilities. Interface feels modern and is relatively easy for fund ops teams to learn. **[Visible](https://visible.vc/)**: Portfolio monitoring: collecting metrics, financials, and qualitative updates from portfolio companies as an ongoing data pipeline. Key features are founder-friendly collection workflows and a flexible data model. Good fit if you already have capital calls and LP financials covered in another tool. ## Things to Consider When Evaluating Beyond core functionality, there are a few things worth thinking about when choosing a platform. **Ecosystem lock-in**: Some platforms, particularly Carta, want you in their full ecosystem: fund administration, cap table management, portfolio company equity. This can be convenient if you're already committed, but makes switching costs high. If you use a different fund administrator or want flexibility, make sure the platform works well standalone. **Data model flexibility**: Fund operations involves messy data. Companies have multiple funding rounds, complex cap tables, and valuations that change quarterly. Some platforms are rigid. Fixed fields, fixed workflows, fixed reports. Others let you customize. The tradeoff is that rigid platforms are faster to set up but frustrating when your process doesn't match their assumptions. Ask how the platform handles edge cases specific to your fund. **Cross-team usability**: Fund ops software often ends up siloed in the finance team, even when investment teams would benefit from access. If your goal is firm-wide visibility into portfolio performance, evaluate whether the interface is intuitive enough for non-finance users. A platform your CFO tolerates but your investment team ignores creates knowledge silos. **Fund modeling realism**: Fund forecasting tools let you model scenarios: deployment pace, follow-on reserves, exit timing. These can be useful for LP communication and planning. But be realistic about their limits. VC returns are driven by outliers, not averages. A model built on "average conversion rates" may feel rigorous but miss how the industry actually works. Use forecasting for communication and planning, not as a source of truth about future performance. **Founder experience**: If your platform collects data directly from portfolio companies (financials, KPIs, updates), consider the founder experience. Multiple investors using different collection tools means founders get bombarded with redundant requests. Some platforms are more founder-friendly than others. This affects response rates and data quality. ## Making the Transition The hardest part of adopting fund operations software isn't choosing the platform. It's transitioning your data and getting your fund ops team to actually use it. **Onboard your fund ops team first**: Don't surprise your CFO or fund administrator with a new system. Involve them in the decision. Have them test the platforms. Get their buy-in. They need to use this regularly. If they find it confusing or harder than their current process, they'll resist and you'll end up maintaining two systems. The best implementations start with "what will make this easier for the operations team?" not "what has the most features?" A system they actually use is better than a sophisticated system they work around. **Transition quickly**: The worst situation is maintaining two systems of record for your fund's finances. Dual entry creates errors, inconsistency, and extra work. Set a hard cutoff date. Move all data into the new system. Deprecate the old system. Don't let "we'll transition gradually" drag on for months. Plan the transition during a quiet period if possible. After you've closed quarterly LP reporting but before the next quarter starts. This gives you time to import historical data and verify everything is correct before you need to produce reports. **Integration with existing providers**: If you outsource some fund operations (to Carta fund administration, for example, or a third-party fund administrator), make sure your operations platform integrates with them. You don't want to manually enter data from your fund admin into your operations platform. The systems should sync or have clean data import/export. **Start with the basics**: Don't try to configure every feature on day one. Get capital management, portfolio tracking, and LP reporting working first. You can add fund forecasting models, portfolio metrics tracking, and sophisticated analytics later. The core operations need to work before you optimize. ## Beyond Operations: Fundraising Value Fund operations software isn't just for managing your current fund. It becomes valuable when you're raising your next fund. **Track record materials**: When you raise Fund II, you need to show Fund I's track record. Portfolio company performance, investment pace, realized returns, DPI and TVPI metrics. A good operations platform can generate these materials in formats LPs expect. You're not building presentation materials from scratch or manually pulling data from spreadsheets. **Portfolio summaries**: LPs want to see your portfolio companies, investment dates, check sizes, current valuations, and performance. Operations platforms can generate these summaries automatically with current data. As portfolio companies raise new rounds or exit, the system updates. Your fundraising materials stay current. **Performance metrics**: IRR, TVPI, DPI, MOIC calculated consistently and presented in standard formats. LPs compare these metrics across many funds. Having clean, auditable data from your operations platform makes fundraising more credible than metrics calculated in spreadsheets. This isn't the primary reason to implement fund operations software. But when you're raising your next fund, you'll be grateful you have clean historical data and can generate track record materials quickly. ## The Bottom Line Buy fund operations software. Don't build it. This is financial infrastructure that needs to be correct, auditable, and maintained. Choose based on integration with your existing systems and what your fund ops team finds easiest to use. The exception: very large funds with significant engineering and operations resources have built custom fund operations platforms including LP portals, automated reporting, and investor relations tools. If you're managing multiple billion-dollar funds with dedicated platform teams, custom infrastructure might make sense. For everyone else, even large established funds, buying off the shelf is the right choice. This is not the most exciting part of building a VC tech stack. But it's essential infrastructure that every fund needs. Get it set up correctly and let it run in the background while you focus on finding and supporting great companies. In the next chapter, we'll look at portfolio support tools, which help you add value to your portfolio companies after you invest. # Fundraising Source: https://buildingfor.vc/guide/part-2-tech-stack/fundraising How to use technology to support raising your next fund without making the process feel impersonal. ## Overview Fundraising is how you raise your next fund. After you've invested Fund I and shown results, you need to convince LPs to commit capital to Fund II. This involves identifying potential LPs, preparing materials that demonstrate your track record, managing relationships over months of conversations, and closing commitments. Most of the infrastructure you need for fundraising already exists in tools we've covered in previous chapters. Fund operations software generates your track record materials. Your CRM tracks LP relationships. Sourcing-type tools can help identify potential LPs. Unlike other parts of the VC tech stack, there's no dominant fundraising-specific platform yet. That could be an opportunity. Fundraising is fundamentally relationship-driven, LPs are making long-term commitments based on trust, but there's room for better tooling that helps funds manage the process more systematically while maintaining the personal touch. This chapter covers what fundraising actually involves from a technology perspective and what tools you already have that support it. ## The Fundraising Process and Tools From a technology and data perspective, fundraising has a few key components. Most of what you need is covered by tools you already have. **Identifying potential LPs**: Before you can raise money, you need to know who to talk to. This is like sourcing, but for LPs instead of companies. You want to identify institutional investors, family offices, high-net-worth individuals, and fund-of-funds that invest in your stage, geography, and strategy. Use the same sourcing tools you'd use for finding companies: [PitchBook](https://pitchbook.com/) has good data on institutional investors, [Specter](https://www.tryspecter.com/) can help identify active LPs in your space. This is less systematic than company sourcing because LP lists are smaller and more relationship-driven, but the same research principles apply. If you're interested in specific datasets for LPs, [Preqin](https://www.preqin.com/) is more focused on LP data. **Finding events and building relationships**: Much of fundraising happens at conferences, LP events, and industry gatherings. Track which LPs attend which events, what conferences matter for your stage and geography, and where you can create in-person touchpoints. Some funds maintain spreadsheets of relevant events with expected LP attendance. The same CRM and research approach that helps you identify LPs can help you understand where to meet them. There are some data providers focused on events (e.g., [Sourcescrub](https://www.sourcescrub.com/)) which can help find the right gatherings for fundraising. **Preparing materials**: LPs want to see your track record, fund strategy, team backgrounds, and portfolio performance. Your [fund operations platform](/guide/part-2-tech-stack/fund-operations) generates performance data and portfolio summaries automatically. Export this into your pitch deck and fundraising materials. Don't build separate systems for fundraising collateral, it should come directly from your operations data. **Managing the process**: Fundraising takes months. You're talking to dozens of potential LPs at different stages: some in early conversations, others doing diligence, some committed. Use your CRM to track LP prospects and conversations exactly like deal flow. LPs move through stages (initial contact, meeting, diligence, committed). Track what materials they've seen, what questions they've asked, and when you need to follow up. Some funds maintain separate systems for LP relationships versus company relationships, but this usually isn't necessary unless you have very different workflows or dedicated teams managing them. **Data rooms**: When LPs do diligence, they want access to detailed fund information: legal documents, portfolio details, historical performance, compliance materials. [Docsend](https://www.docsend.com/) is the standard. Organize your materials into a clean structure, grant access to LPs as they enter diligence, and track what they're looking at to understand their concerns and interests. Docsend analytics show you which documents get the most attention, which helps you prepare for their questions. **Ongoing communication**: Throughout fundraising, the GPs and investor relations team will be sending updates, answering questions, and staying in touch with prospects through email and meetings. Use your CRM to track conversations so you maintain context over months of fundraising. LPs often ask similar questions across diligence processes, maintain a document with standard questions and your answers, with links to relevant documents in your data room. This helps you be responsive without making it feel automated. **The key insight**: Focus on internal tools that help you be more organized, responsive, and prepared. Better CRM so you never forget an LP conversation. Better operations software so you can generate track record materials quickly. Good data room organization so LPs can find what they need during diligence. The tools exist - use them well. ## The Bottom Line Fundraising is covered by tools you already have. Use sourcing tools to identify potential LPs, fund operations software to generate track record materials, your CRM to manage relationships, and Docsend (or equivalent) for data rooms. Focus your engineering time on tools that improve your investment process (research, deal flow, operations). Fundraising happens episodically (every few years) and is relationship-driven. The tools you already have are sufficient if you use them well. In the next chapter, we'll look at your fund's website and public presence, which is often simpler than you think but still important to get right. # Introduction to the VC Tech Stack Source: https://buildingfor.vc/guide/part-2-tech-stack/introduction Understanding what software exists at VC funds, who cares about what, and how to prioritize building the right things. ## Overview You've learned the VC fundamentals, understood your fund, and know the common mistakes to avoid. Now it's time to understand what you might actually build. Part 2 covers the different types of software that exist at VC funds, what they do, and when to prioritize building them. Before diving into each category, it's important to understand that different parts of the tech stack matter to different people at the fund. Not every tool serves every stakeholder. Understanding who cares about what helps you find the right team to collaborate with on building your product. ## The Technology Landscape Picture a typical scenario at most funds. Research is fragmented. Notes live in people's heads, scattered across email threads, Slack messages, Google Docs, Notion pages, and personal notebooks. When someone researched a similar company six months ago, you can't find their work. When you're evaluating a founder who pitched you two years ago, you can't remember what you learned back then. Research gets done, but it doesn't accumulate into institutional knowledge. Deal flow is managed in spreadsheets or generic CRMs that don't quite fit the workflow. Portfolio companies report metrics through email and inconsistent formats. LP reporting is a quarterly scramble to collect data from multiple sources. The fund's website is static and rarely updated. Fundraising materials are cobbled together from various documents. This isn't because funds don't understand technology. It's because VC tools are niche, workflows are specific, and off-the-shelf software rarely fits perfectly. Building custom tools can be transformative, but only if you build the right things for the right people at the right time. ## Who Cares About What Different parts of the tech stack serve different stakeholders: ### [Research Platforms](/guide/part-2-tech-stack/research-platforms) **Primary users**: GPs, Partners, Associates doing research **Value**: Clarity of thought, thesis development, competitive advantage ### [Sourcing Tools](/guide/part-2-tech-stack/sourcing-tools) **Primary users**: Deal team, Partners doing outbound **Value**: Finding companies proactively, not just reacting to inbound ### [CRM / Deal Flow Management](/guide/part-2-tech-stack/crm-and-deal-flow) **Primary users**: Everyone on the investment team **Value**: Tracking relationships, managing pipeline, coordinating team ### [Fund Operations](/guide/part-2-tech-stack/fund-operations) **Primary users**: CFO, Fund Administrator, Operations team **Value**: Capital calls, distributions, LP reporting, compliance ### [Portfolio Support](/guide/part-2-tech-stack/portfolio-support) **Primary users**: Platform team, Operating Partners, Deal Team **Value**: Helping portfolio companies succeed, value-add services ### [Fundraising](/guide/part-2-tech-stack/fundraising) **Primary users**: Managing Partner, Investor Relations **Value**: LP relationship management, fundraising materials ### [Website & Public Presence](/guide/part-2-tech-stack/website-and-external-presence) **Primary users**: Marketing, Investor Relations **Value**: Brand, inbound deal flow, LP communication ## Understanding Priorities Not all of these matter equally at every fund. A small seed fund with two partners has very different needs than a growth equity fund with 30 employees. Your priorities depend on: **Fund size and stage**: Early-stage funds need research and deal flow tools. Later-stage funds need robust operations and portfolio management. Multi-stage funds need everything. **Team structure**: Solo GPs can manage with simpler tools. Teams of five or more need collaboration infrastructure. Funds with dedicated platform teams need portfolio support tools. **Strategy**: Thesis-driven funds need research platforms. Network-driven funds need CRM. Operator funds need portfolio support tools. **Existing tools**: If you already have Affinity for CRM and Carta for fund admin, don't rebuild those. Build what's missing or what's uniquely valuable to your strategy. **Technical resources**: One person can realistically build and maintain 1-2 substantial tools. Choose carefully. ## How This Section Works Each chapter in Part 2 covers one category of VC software: * What it does and why it matters * Key features and workflows * Build vs. buy considerations * When to prioritize it * Common tools and approaches * Real examples where applicable By the end of Part 2, you'll understand the full landscape of what you could build and have a framework for deciding what to prioritize for your specific fund. We start with research platforms because they're the foundation. Before you can source deals, manage pipeline, or support portfolio companies, you need clarity about where to invest. Research creates that clarity. # Portfolio Support Source: https://buildingfor.vc/guide/part-2-tech-stack/portfolio-support Understanding what portfolio support actually looks like, what can and can't be automated, and why strategy matters before building tools. ## Overview Portfolio support is where many VCs want to differentiate. The pitch is compelling: we don't just write checks, we help our portfolio companies succeed. We provide resources, connections, and tools that make founders more likely to build successful companies. Technology and data can help with some of this, but portfolio support is fundamentally different from the other parts of the VC tech stack we've covered. Research platforms, sourcing tools, CRMs, and fund operations are infrastructure that supports your investment process. Portfolio support tools are meant to help your founders run their companies better. The challenge is that most portfolio support doesn't fit into repeatable products. Every company has different needs. What helps one founder find customers might be irrelevant to another. The patterns that justify building software (doing the same thing repeatedly, at scale, with clear value) often don't apply to portfolio support. This chapter covers what portfolio support actually looks like, what can and can't be automated, and why you should start with small projects before building platforms. More importantly, it covers why you need to figure out your strategy around portfolio support before committing engineering resources. ## What VCs Actually Help With When VCs talk about portfolio support, they usually mean a handful of core areas with different levels of tech leverage: **Talent** - Helping portfolio companies hire key people. Technology helps: tracking talent networks, surfacing relevant candidates, managing introduction workflows. Human value: your personal network and ability to sell candidates on the opportunity. **Quick win: Portfolio jobs board** One of the easiest ways to provide immediate talent value to your portfolio is setting up a jobs board. Tools like [Getro](https://www.getro.com/) or [Consider](https://consider.com/) automate this entirely: they aggregate job postings from your portfolio companies and create a branded careers page for your fund. No engineering work required, and it gives founders a recruiting channel on day one. **Fundraising** - Supporting portfolio companies raising their next round. Technology helps: maintaining investor relationship data, tracking which investors invest in which spaces. Human value: relationships with other investors and judgment about fundraising strategy. **Customer discovery** - Helping portfolio companies identify potential customers, especially B2B. Technology helps: using datasets (PitchBook, LinkedIn) to identify ideal customer profiles and generate targeted prospect lists. Human value: making warm introductions and coaching on sales strategy. **Exit planning** - Helping companies think about acquisitions or IPO readiness. Technology helps: tracking potential acquirers, understanding market dynamics. Human value: making introductions to bankers or acquirers, judging timing and strategy. **Board participation and strategic guidance** - Strategic advice, product feedback, go-to-market strategy. Technology helps: automated note-taking tools ([Granola](https://www.granola.ai/), [Otter](https://otter.ai/)) that capture discussion points and sync to your CRM, reducing cognitive overhead during meetings so partners can focus on the conversation. Human value: deep context about the company, market understanding, and judgment from experience. This is where being a good board member matters. The pattern: technology enhances the connection-making and information-organizing parts of portfolio support. It can't replace the relationships, judgment, and specific expertise that make portfolio support actually valuable. ## The Project vs. Product Problem Most portfolio support work doesn't fit into repeatable products. This is the fundamental challenge that makes it different from other parts of your tech stack. **Why products work elsewhere**: A CRM is a product because every fund tracks relationships and deals the same way. Fund operations software is a product because every fund needs to manage capital calls and LP reporting. The workflows are similar enough across users that building a product makes sense. **Why portfolio support is different**: Every company has different needs at different times. A deep tech hardware company needs help hiring mechanical engineers and navigating government contracts. A consumer social app needs help with growth marketing and app store optimization. A B2B SaaS company needs help with enterprise sales. There's no common workflow or repeatable pattern. You can't build "the portfolio support platform" because there's no single thing that all portfolio companies need. What you end up with are projects: one-off tools or analyses that help specific companies with specific problems. **What this means in practice**: When a portfolio company asks for help finding enterprise customers, you might build a quick analysis of potential customers in their space. When another company is hiring a CTO, you might pull together data on relevant candidates from your network. These are valuable, but they're projects, not products. The next company will have different needs. The implication is that you should be very careful about investing significant engineering time in portfolio support platforms. You'll likely build something that helps 2-3 companies but doesn't generalize. Meanwhile, you could have been improving your research platform or deal flow tools that help your entire investment process. ## Start Small and Validate The biggest mistake funds make with portfolio support is building platforms before validating that founders actually want them. **The pattern**: A fund decides they want to differentiate on portfolio support. They imagine a platform where founders can find candidates, get customer intros, access market data, connect with each other. They spend six months building it. Founders don't use it. The reasons vary: it doesn't solve a painful enough problem, founders already have other solutions, the UX isn't good enough to displace existing tools, or the value proposition isn't clear. You've wasted engineering time. More importantly, you've wasted founder time asking them to test and give feedback on something that doesn't help them. **Start with projects instead**: When a founder asks for help, solve that specific problem. Build a lightweight tool or analysis. See if they use it. See if other founders have the same need. If you do the same type of project 5-10 times, then maybe it's worth building a product. More likely, you'll end up with a collection of scripts and templates that speed up future projects without becoming a maintained product. This is the opposite of how you'd approach building internal tools. For research platforms or deal flow management, you can design the system upfront because you understand your own workflow. For portfolio support, you're building for your founders' workflows, which are diverse and changing. You need validation before committing to products. ## Different Technology Strategies for Portfolio Support Funds approach technology and data for portfolio support very differently. There's no consensus on what the right strategy is. Understanding the range of approaches helps you figure out where you want to be. **No technology investment**: Many funds, including large ones, explicitly don't invest engineering or data resources in portfolio support. The reasoning is that it doesn't scale and they have limited resources. They'd rather focus their data and engineering teams on research platforms, deal flow tools, and fund operations, things that directly improve their investment process. This is a legitimate strategy that lets you focus limited technical resources where they create the most leverage. **Opportunistic projects**: Help with specific data or technology projects when founders ask, but don't build systematic infrastructure. If a portfolio company needs customer discovery help, you might pull together a prospect list. If another needs market analysis, you might do that analysis. But you're not building platforms or repeatable tools. This is where most funds with data teams end up. You use your technical capabilities to help when it makes sense, but you don't commit to systematic portfolio support. **Strategic technology investment**: Focus on specific high-value areas where technology creates real leverage. Maybe you build tools for customer discovery because many of your B2B portfolio companies need that. Or you build recruiting workflows because you invest heavily in helping portfolio companies hire. This is selective: you identify areas where technology helps repeatedly and invest there, but you don't try to build comprehensive portfolio support infrastructure. **Platform-driven support**: Make portfolio support a core part of your value proposition and invest significant resources (people and technology) in it. [YC](https://www.ycombinator.com/) is the extreme version: their portfolio support infrastructure ([Bookface](https://bookface.ycombinator.com/) for founder connections, [Work at a Startup](https://www.workatastartup.com/) for talent, [Startup School](https://www.startupschool.org/) for founder best practices) is as important as their capital. This requires real commitment: dedicated platform teams, engineering resources, and making it central to your strategy. Only makes sense if you have the scale, community, and strategic focus to make it work. ## Figure Out Your Strategy First Before you build anything for portfolio support, be explicit about your strategy. What will you do? What won't you do? Why? **Questions to answer**: * Is portfolio support core to our value proposition or supplementary? * Do we have the resources (people, time, network) to deliver meaningful support? * What specific types of support can we actually provide better than founders could get elsewhere? * Will data and technology genuinely make our support more valuable, or are we building because it sounds good? * If our engineering time is limited, will we get more value from portfolio support or from improving our internal tools? ## Real Example: What Actually Gets Used **Author Note: Portfolio Support Projects** At Inflection, we focused portfolio support on two areas: hiring support and customer discovery. When portfolio companies needed help finding talent or when B2B companies needed to identify potential customers, we could pull together prospect lists from datasets. What's notable is that these have all been projects, not products. We've never built a repeatable tool that multiple portfolio companies use regularly. Every request is different, that's great, it's fun to work with the founders. ## The Bottom Line Portfolio support is fundamentally different from other parts of the VC tech stack. Most of it doesn't fit into repeatable products. Every company has different needs at different times. Start with projects, not platforms. When a founder asks for help, solve that specific problem. If you solve the same problem repeatedly, then consider whether a product makes sense. But don't build portfolio support infrastructure before validating that founders actually need it. You'll waste your time and theirs. More importantly, figure out your strategy first. Some funds don't do systematic portfolio support at all, and that's fine. Others make it central to their value proposition with dedicated teams and resources. Most are somewhere in the middle. Be explicit about where you are and align your engineering resources accordingly. If your goal is to build data and technology infrastructure for your fund, you'll likely get more value from improving your research, deal flow, or operations tools than from building portfolio support platforms. Portfolio support is where human judgment, relationships, and specific expertise matter most. Technology can enhance these things at the margins, but it can't replace them. In the next chapter, we'll look at fundraising tools, which help you raise your next fund and manage LP relationships. # Putting It Together Source: https://buildingfor.vc/guide/part-2-tech-stack/putting-it-together How to sequence your tech stack implementation across three phases - learning, building the backbone, and taking your big swing. ## Overview You've now seen seven chapters covering different categories of tools and technology that venture funds use: research platforms, sourcing tools, CRM and deal flow management, fund operations, portfolio support, fundraising infrastructure, websites, and external presence. Each has value. Each can improve how your fund operates. You don't need all of them immediately. The biggest mistake new technical hires at VC funds make is trying to build everything at once. They see the full landscape of what's possible and want to implement a comprehensive tech stack from day one. This spreads resources too thin and delivers less value than focusing on fewer things done well. This chapter is about sequencing. There are three phases: learn what your fund actually needs, build the backbone and quick wins, then decide on your big swing. Don't skip phases. ```mermaid theme={null} flowchart LR P1[Phase 1: Learn] --> P2[Phase 2: Backbone + Quick Wins] --> P3[Phase 3: Big Swing] ``` ## Phase 1: Learn What Your Fund Needs First Before you build anything significant, spend time [understanding your fund](/guide/part-1-understanding-vc/understanding-your-vc-fund). Observation comes before building. **Talk to the people doing the work**: Meet with GPs and partners about how they find companies, make decisions, and track portfolio companies. Talk to the fund operations team about LP reporting and capital management. Ask what's painful right now, not what would be cool to have. **Understand the thesis and strategy**: Is your fund thesis-driven or network-driven? Are you pre-seed or growth equity? Do you specialize in a specific vertical or geography? The answers dramatically change what technology matters. A thesis-driven fund needs research infrastructure. A network-driven fund needs relationship tracking. Don't assume what works at other funds works at yours. **Identify the actual pain points**: Watch for what's actually broken or painful, not what people say would be nice to have. Is deal flow tracking chaotic? Is research getting lost? Are portfolio companies asking for help you can't deliver systematically? These are signals about where to invest. **Don't assume**: Just because EQT built Motherbrain or Inflection built Kepler doesn't mean your fund needs the same thing. Every fund is different. Some funds operate perfectly well with minimal technology. Others gain competitive advantage through sophisticated infrastructure. Figure out which you are before committing to major projects. This learning phase might take a few months. That's fine. Building the wrong thing wastes more time than taking time to understand what to build. ## Phase 2: Build the Backbone and Quick Wins Once you understand your fund, start with two parallel tracks: the backbone (foundational infrastructure everyone needs) and quick wins (small tools that prove value). **The backbone: CRM and fund operations** These are non-negotiable infrastructure. Every fund needs to track deal flow and manage fund operations. Don't build these. Buy them. **CRM**: See [CRM and Deal Flow](/guide/part-2-tech-stack/crm-and-deal-flow). Get it set up immediately and make sure your team actually uses it. If adoption is low, you have a process problem to solve before building anything else. **Fund operations**: See [Fund Operations](/guide/part-2-tech-stack/fund-operations). If you're Fund I with a small LP base, spreadsheets might work temporarily. But plan the migration before Fund II. Buy this off the shelf. Don't build fund operations software. These aren't exciting projects, but they're essential. Get them working before anything else. **Quick wins: small tools that build trust** While implementing the backbone, build small tools that solve specific pain points and prove your value to the team. These establish credibility and help you understand what resonates. Examples of quick wins: * A simple dashboard showing portfolio company metrics pulled from your CRM * Automating a painful manual process (generating LP reports, tracking follow-on opportunities) * A lightweight research workflow improvement (better way to organize market maps, easier way to share insights) * Customer discovery prospect lists for portfolio companies (using data from PitchBook or LinkedIn) Quick wins should take days or weeks, not months. They should solve real problems people have right now. They should be immediately useful without requiring behavior change from the team. These projects do three things. First, they deliver tangible value. Second, they build trust with the investment team. Third, they help you test out new technology. When you later propose a bigger project, you've proven you can deliver useful things rather than over-engineered solutions nobody uses. ## Phase 3: Decide on Your Big Swing After you've built the backbone and delivered quick wins, you've earned the credibility and understanding to take on one big project. This is where you differentiate. You get one big swing. Maybe it's a research platform. Maybe it's sophisticated sourcing infrastructure. Maybe it's portfolio support tools. Choose based on where your fund's competitive advantage actually is, not what sounds impressive. And maybe you don't need a big swing at all. Many successful funds operate with minimal technology beyond the backbone. Don't build because you think you should. **Align on data needs and budget first** Before committing to your big project, align with fund leadership on what data you'll need and ensure you have proper funding for both the data and the build. Data subscriptions are expensive. PitchBook costs tens of thousands per year. Harmonic and Specter are significant investments. LinkedIn Sales Navigator for your whole team adds up. If your big swing is a research platform that aggregates market data, or sourcing infrastructure that needs company datasets, or portfolio analytics that requires financial data, you need budget for the underlying data sources. Have this conversation before you start: What data sources will this project need? What do they cost? Do we have budget for them? Who approves data spending? This also applies to engineering resources. If your big swing requires custom development, you need ongoing engineering time to maintain and improve it. Make sure leadership understands this is a strategic investment that requires sustained resources. ## The Bottom Line Three phases. Don't skip them. **Phase 1**: Learn what your fund needs. Spend time observing, asking questions, and understanding the strategy before building anything significant. **Phase 2**: Build the backbone (CRM, fund ops) and quick wins (small tools that solve immediate pain points and prove value). Get the essentials working and establish credibility. **Phase 3**: Decide on your big swing based on where your fund's competitive advantage is. Research platform? Sourcing? Portfolio support? Make it count. This approach avoids [common mistakes](/guide/part-1-understanding-vc/common-mistakes). You're not jumping in too early because you spent Phase 1 learning. You're not over-engineering because Phase 2 focuses on small, proven wins. You're not building in a vacuum because you've validated what matters before Phase 3. Most importantly, you're building momentum. Quick wins in Phase 2 create energy and trust. By the time you're ready for your big swing in Phase 3, you have the credibility and understanding to do it right. In Part 3, we'll go deep on the technical foundations underlying all these tools: data modeling, entity resolution, data quality, warehousing, and integration patterns. These are the building blocks you need to implement any of the tools covered in Part 2. # Research Platforms Source: https://buildingfor.vc/guide/part-2-tech-stack/research-platforms How research platforms create competitive advantage through clarity of thought, thesis development, and institutional knowledge. ## Overview The best advantage you have as an investor is clarity of thought. When everyone else is chasing the same hot deals, clarity lets you see opportunities others miss. When markets shift, clarity lets you adapt your thesis instead of following the herd. Research platforms are the tool that helps you develop and maintain that clarity. Unlike CRM which tracks relationships, research platforms help you think clearly about where to invest in the first place. This chapter covers how research platforms create competitive advantage and what you should build. ## Research as Competitive Advantage Most funds operate reactively. Deals come in through the network. Partners evaluate them. Some get funded, most don't. The fund is a filter: deals flow in, a few flow out with investment. This is how venture capital has worked for decades. The best funds operate proactively. They develop strong points of view about markets, technologies, and trends before seeing deals in those areas. They publish their thinking. They reach out to founders building in spaces they find interesting. When a company in their thesis area raises a round, they're already at the table because founders know their perspective. This approach requires deep research about markets, technologies, business models, and secular trends. Not research about individual companies, that comes later, but macro research that identifies opportunities and develops into investment theses. The problem is that this kind of research is incredibly hard to do well without infrastructure. Where do you store evolving thinking? How do you connect research across different areas? How do you turn scattered insights into coherent theses? Research platforms solve this. ## What Research Platforms Do Research platforms support the full cycle from macro research to investment conviction: **Macro research and thesis development**: Start with broad questions. What's happening in infrastructure software? How is AI changing developer workflows? What new business models are emerging in healthcare? Research platforms give you space to explore these questions, capture insights, connect related ideas, and gradually develop coherent theses. Unlike scattered documents, a research platform shows you connections. Your research about developer tools connects to your research about AI. Your analysis of a specific market connects to broader technology trends. Over time, patterns emerge. These patterns become theses. **Publishing and demonstrating your thinking**: Once you've developed a thesis, you want to share it. Write it up. Publish it on your website or as a report. A research platform becomes the source material for published content. All your research about a topic is already organized and connected. You're not starting from scratch when you want to write something. You're synthesizing research you've already done. Publishing serves multiple purposes. It demonstrates to founders that you understand their space. It attracts inbound from companies building in areas you find interesting. It forces you to clarify your own thinking. The act of writing for an audience sharpens fuzzy ideas into clear arguments. **Building conviction before deals appear**: When you've spent months researching a market, you develop deep conviction. You understand the key players, the technology trends, the business model challenges, and the opportunities. When a company in that space raises a round, you don't need three months of diligence. You already have context. You can move quickly because you've already done the work. Everyone else is learning about the space from the company's pitch. You're evaluating whether this specific company is the right one to back in a space you already understand deeply. **Institutional knowledge and collaboration**: Research compounds. When multiple people on your team research adjacent areas, their insights should connect. One partner researches infrastructure, another researches developer tools, a third researches AI. A good research platform surfaces the connections. The infrastructure research informs the developer tools research. The AI research connects to both. This is how small teams can compete with large ones. Not by having more people, but by making each person's research more valuable through better connections and institutional memory. ## Key Features and Workflows Research platforms that support thesis-driven investing need different features than simple note-taking tools: **Networked research**: Ideas connect to other ideas. A research note about infrastructure connects to notes about developer tools, which connect to notes about specific technologies. The platform shows these connections explicitly. You're not just creating isolated documents. You're building a knowledge graph where insights emerge from the connections. This is where tools like Notion fall short. They're great at hierarchical organization (folders and pages) but weak at networked thinking. You want bidirectional links, automatic backlinks, and the ability to see how different pieces of research relate. **Market and technology research**: Research about markets, technologies, business models, and trends should be first-class entities, not just tags on companies. When you research infrastructure software, that's a substantial piece of work that should live independently. Companies building in that space link to it, but the market research stands on its own. **Thesis development**: Convert research into theses. A thesis isn't just a collection of notes. It's a coherent argument about why a particular area is interesting. The platform should support drafting theses, getting feedback from the team, refining them, and eventually publishing them. **Publishing workflow**: The path from research to published content should be smooth. Export research into article format. Share drafts with the team. Publish to your website. Track which articles get engagement. The platform bridges internal research and external communication. **Company and founder tracking**: Yes, you still need this, but it's supporting infrastructure for the research, not the main event. When you're deep in infrastructure research and a relevant company appears, you want to quickly capture basic information and link it to your broader research. But the company profile is shallow compared to the market research. **Search and connection surfacing**: When you're researching a new area, the platform should surface relevant existing research. You're looking into healthcare AI? Here's the healthcare research, here's the AI research, here's where those areas have intersected before. Let the platform help you avoid reinventing the wheel and find unexpected connections. ## The Build vs. Buy Decision Research platforms are the strongest candidate for building internally because they're so specific to how your fund thinks. Generic tools work for storing notes, but they struggle to support the full workflow from macro research to published thesis. There's also a lack of VC-specific tools in the market in . **When building makes sense**: Your thesis development process is unique. How you explore markets, connect ideas, develop conviction, and publish thinking doesn't fit generic templates. A deep tech fund researching quantum computing needs to capture different connections than a consumer fund researching creator economy trends. The structure needs to match your intellectual process, not a standard note-taking pattern. Publishing integration is critical to your strategy. If going from research to published articles is core to how you demonstrate expertise and attract founders, you need smooth workflows. Basic collaborative platforms can store drafts but struggle to help you synthesize months of networked research into coherent published content. Custom platforms can build this synthesis directly into the interface. Networked thinking is central to your research approach. If your competitive advantage comes from connecting insights across different research areas, you need more than basic hyperlinks. When AI infrastructure research connects to developer tools research connects to specific technology bets, the relationships themselves become valuable. Custom platforms let you model these relationships explicitly. The value compounds exponentially over time. Every piece of research makes future research more valuable. Every connection discovered feeds into better theses. Every published article attracts better inbound. This is infrastructure that creates lasting competitive advantage. **When existing tools are sufficient**: You're just starting thesis development and want to experiment with networked thinking approaches. Off-the-shelf tools built for this workflow let you explore what works before committing to custom development. Your team is small and collaborative needs are basic. If two or three people are doing research, collaborative platforms with simple linking may be enough. The overhead of custom development outweighs the benefits. Research and publishing aren't core to your competitive advantage. If you invest primarily through network deal flow and don't publish thinking externally, sophisticated research infrastructure may be overbuilding. Building takes significant time. Three to six months for something genuinely useful. Longer to build something that supports your full research-to-publishing workflow. Make sure research and thesis development are actually core to your strategy before committing engineering resources. ## When to Prioritize This Research platforms are most valuable at funds that compete on clarity of thought: **Thesis-driven funds**: If your competitive advantage comes from deep conviction about where technology and markets are going (and making "moonshot bets"), research platforms are core infrastructure. You need to capture how you're thinking about a space, evolve that thinking over months, and eventually publish it to demonstrate your perspective. This is the foundation of your strategy. **Funds that publish thinking**: If you write articles, publish reports, or create content demonstrating your expertise, a research platform becomes your content engine. All your research feeds into published pieces. You're not starting from scratch every time you write. You're synthesizing months of accumulated insights. **Collaborative research teams**: If multiple people research adjacent areas and you want their insights to connect, you need infrastructure for networked thinking. One person researches AI, another researches developer tools, a third researches infrastructure. The connections between their work create unique insights. Don't prioritize research platforms if you invest reactively based on network deal flow, don't develop public theses, are a solo GP with simple note-taking needs, or focus on late-stage deals where diligence is mostly quantitative analysis of business metrics. ## Common Tools and Approaches Most funds should start with existing tools to understand their research workflow before building custom platforms. **Networked thinking tools ([Obsidian](https://obsidian.md/), [Roam](https://roamresearch.com/), [Anytype](https://anytype.io/))**: Built specifically for networked thinking with bidirectional links and graph views. Obsidian is especially popular with individual researchers doing deep thesis work. The markdown-based approach means your content isn't locked in. You can version control your research with Git, use it alongside Claude Code, and export to any format. The main limitation is collaboration. These tools are built for individual use. Team collaboration requires sync solutions (Obsidian Sync, Git repositories, or shared folders) which work but feel less natural than true collaborative platforms. If you're a solo GP or small team comfortable with these workflows, they're excellent starting points. When to use: Starting thesis development, experimenting with networked thinking, individual researchers who want powerful linking, teams comfortable with technical sync solutions. **Collaborative platforms (Notion)**: Flexible and easy for teams to adopt. Many funds use Notion for research because everyone already knows how to use it. Real-time collaboration works well. The weakness is that Notion is hierarchical by design, not network-first. You can link pages, but the mental model is folders and databases, not a knowledge graph. This matters less if your research is more structured and less if you're exploring complex connections between different research areas. Notion works well for collaborative drafts and can export to publishing platforms. When to use: Teams that need easy collaboration, funds with more structured research processes, quick setup without technical complexity, good enough solution while you figure out what's missing. **Custom-built platforms**: Funds serious about research as competitive advantage often build custom. Usually web-based (for Inflection: Next.js frontend, Postgres backend) with strong full-text search (pg\_search or pg\_vector extensions for semantic search). The key technical challenges are more about product design than technology. How do you model relationships between research entities? What makes relevant connections surface naturally? How do you build publishing workflows that feel smooth? How do you ensure that the research experience is as good as your competition: Claude, ChatGPT, and Perplexity? When to use: Research is core competitive advantage, you've outgrown existing tools and know exactly what you need, you have engineering resources, thesis development and publishing are central to your strategy. **Recommended approach**: Start with Obsidian (if comfortable with markdown and sync solutions) or Notion (if team collaboration matters more). Use it for six months. Develop your research process. Figure out how your team actually explores markets, connects ideas, and develops theses. Pay attention to what's frustrating. Can't find connections you know exist? Publishing workflow is too manual? Research isn't connecting to sourcing? These pain points tell you what to build. When you do build custom, you'll know exactly what features matter because you've felt their absence. The key is making the path from internal research to published thinking as smooth as possible. If publishing demonstrates your thinking and attracts founders, the research-to-publishing workflow should be friction-free. This is often the feature that pushes funds toward custom platforms. ## Real Example: Kepler at Inflection **Author Note: Building Kepler** [Kepler](https://svrgn.substack.com/p/introducing-kepler-inflections-home), Inflection's research platform, came from watching how the GPs actually worked. Inflection was deeply thesis-driven. The partners spent significant time researching markets, technologies, and trends before deals showed up. That research was scattered across Notion pages, Signal threads, and individual notes. The problem wasn't organization. It was that insights weren't connecting. One partner's research about the changing European Defense Landscape couldn't easily inform another's work on counter UAV solutions. When it came time to write published content demonstrating our thinking, we'd start from scattered notes instead of synthesized research. Kepler became the platform for thesis development. Market research, technology analysis, and meeting notes all lived together with explicit connections. When you researched AI Agents, you saw related infrastructure research and relevant developer tools work. The platform made it easier to see patterns across different areas of research. The real value came when partners started using it as the source for published thinking. Instead of writing from scratch, they'd synthesize months of research captured in Kepler. The platform went from research storage to thinking tool to content engine. ## The Bottom Line Research platforms are about developing clarity of thought and turning that clarity into competitive advantage. If your fund competes on deep conviction about markets and technologies, research is your foundation. Everything else builds on that. The best investors don't just see more deals. They think more clearly about where to invest. Research platforms help you develop that clarity, maintain it over time, collaborate on it with your team, and demonstrate it publicly to attract the best founders. Don't build a research platform to organize notes better. Build it to think more clearly, publish more coherently, and move faster when opportunities appear in spaces you already understand deeply. In the next chapter, we'll look at sourcing tools, which help you find companies in the spaces your research has identified as interesting. # Sourcing Tools Source: https://buildingfor.vc/guide/part-2-tech-stack/sourcing-tools Understanding when automated sourcing actually creates value ## Overview Sourcing is where every fund wants to start when they bring in their first data person. The pitch is compelling: use data to systematically find companies before your competition, build relationships early, get into deals proactively rather than reactively. But sourcing tools don't work equally well for all funds, and understanding when they create value matters more than understanding how to build them. This chapter covers the history of sourcing infrastructure at VC funds, what sourcing tools do, when they actually create value, and why you should probably buy rather than build unless you've found a very specific niche where custom tooling creates real advantage. ## The Historical Advantage Around 2017-2018, funds that invested in sourcing infrastructure got real competitive advantage. EQT Ventures built Motherbrain. SignalFire created their talent network and technical signal tracking. Early Bird and Moonfire developed sophisticated data operations. These funds could systematically find companies others missed, reach out early, and build relationships before competitive fundraising processes. This worked because the infrastructure didn't exist yet. There were no purpose-built VC sourcing tools. Data providers had limited coverage. If you wanted systematic sourcing, you had to build it yourself. The funds that invested in this infrastructure got ahead and proved the model worked. Their success changed the market. The existence of tools like Harmonic and Specter is a direct result of these pioneers demonstrating that systematic sourcing creates value. Data providers improved their coverage because funds demanded it. The baseline rose industry-wide. These early movers still maintain advantages: proprietary data from years of tracking, refined processes, established relationships from early outreach, and continued innovation beyond what off-the-shelf tools provide. But the barrier to entry for basic sourcing capability has collapsed. ## What Sourcing Tools Do Sourcing tools help you discover companies, track them over time, and manage outreach systematically. **Discovery**: Find companies that match your criteria. You're interested in developer tools focused on AI infrastructure? You want to see every company in that space. Sourcing tools aggregate data from multiple sources: web scraping, funding databases, job postings, social media, technical signals. They let you search and filter to identify relevant companies. Different funds care about different signals. Some look for companies with strong technical teams (GitHub activity, Stack Overflow contributions). Others care about traction (job growth, web traffic, social media presence). Some focus on geography or funding stage. The key is turning your investment criteria into searchable filters. **Tracking**: Once you've identified interesting companies, you need to track them over time. Is their team growing? Did they just raise a round? Are they hiring for key roles? Good sourcing tools monitor companies automatically and surface changes (signals) worth paying attention to. This is where many manual processes break down. You find 50 interesting companies. Six months later, you can't remember which ones you were tracking or why. Sourcing tools maintain that institutional memory and alert you when something changes. **Enrichment**: Basic information isn't enough. You want context. Who are the founders? Where did they work before? Who's already invested? What technologies are they using? How much traction do they have? Enrichment layers on data from multiple sources to give you a fuller picture. The challenge is that no single data source has everything. LinkedIn has professional backgrounds. Crunchbase has funding data. GitHub shows technical activity. Good sourcing tools combine multiple data sources to create a more complete view. **Outreach**: Identifying companies is only half the work. You need to reach out. Sourcing tools help manage that process. Track who you've contacted. Follow up systematically. Measure response rates. Some tools integrate email sequencing, others just track outreach in your CRM. The key is making proactive sourcing systematic rather than sporadic. You're not just searching once. You're continuously monitoring spaces you care about and building relationships with interesting companies over time. ## When Sourcing Actually Works Before deciding whether to build or buy sourcing tools, you should question whether you need them at all. Sourcing tools work well for some funds and poorly for others. The determining factors are stage, strategy, and specificity. **Stage matters more than most people think**. Sourcing tools rely on companies existing in datasets. If you're investing at Series A or later, companies have usually raised previous rounds, hired employees, built websites, and generated signals that data providers capture. You can find them. But if you're investing pre-seed or at the earliest stage of seed, companies often don't exist in any dataset yet. They're two founders working nights and weekends. No website, no job postings, no press releases, no funding data. Sourcing tools can't find what isn't there. More fundamentally, pre-seed investment decisions aren't based on data you can source. You're investing in the founders' resilience, domain expertise, ability to execute, and judgment. These qualities can't be inferred from a LinkedIn profile or GitHub activity. You need conversations. Your deal flow comes from network and research, not systematic sourcing. **Thesis-driven investing creates a signal-to-noise problem.** If your criteria are relatively broad (B2B SaaS companies in the US with 10-50 employees), filters work well. But thesis-driven funds with specific criteria (Hardware startups building autonomous vehicles for mining) get overwhelmed with noise. You get hundreds of companies that match broad filters but aren't actually relevant to your specific interest. Sorting through them takes more time than the sourcing saves. Your competitive advantage comes from understanding spaces better than others, and that understanding comes from research, not filtering databases. **Network-driven funds might not need sourcing at all.** If your competitive advantage is relationships that lead to warm introductions, you're not trying to systematically find companies. Sourcing tools don't match your workflow. **The uncomfortable truth**: Many early-stage, thesis-driven funds think they need sourcing tools because it sounds like the right thing to do. But their companies are too early for datasets, their criteria are too specific for filters, and their competitive advantage comes from research and relationships, not systematic discovery. Sourcing tools become shelf-ware. ## The Build vs. Buy Decision Don't build generic sourcing infrastructure. The tools exist (Harmonic, Specter, amongst others), they work well, and the data is commoditized. Building this in is like building your own CRM instead of using something off-the-shelf. Harmonic and Specter aren't perfect. No sourcing tool will be. But they're good enough that rebuilding their entire platform internally doesn't make sense. You'll spend months recreating features they already have, and you'll end up with something similar but worse. Use that engineering time for things that actually differentiate your fund. **The exception**: genuinely unique niches where custom tooling creates real advantage. Some funds have built knowledge graphs tracking interactions to optimize warm intro paths. Others have built deep open source analysis tools (GitHub stars, commit velocity, code quality, contributor growth) for investing exclusively in open source companies. These work because they're truly specialized, not just "we want different filters." **The realistic recommendation**: Series A+ with broad criteria? Buy Harmonic or Specter. Pre-seed/early seed with specific criteria? You probably don't need sourcing tools at all. Genuinely unique niche requiring specialized analysis? Consider building that specific piece. ## Common Tools **Harmonic**: The most popular purpose-built VC sourcing platform. Aggregates data from multiple sources, supports search and filtering, tracks companies over time, integrates with CRMs. **Specter**: Another purpose-built option with emphasis on technical signals and growth indicators. Similar feature set to Harmonic with slightly different data sources. **Crunchbase or PitchBook**: Traditional funding databases. Useful for basic research and market mapping but not sophisticated enough for proactive sourcing on their own. See [Signal Data Providers](/guide/part-3-technical-foundations/data-providers/signal-data) and [Company Data Providers](/guide/part-3-technical-foundations/data-providers/company-data) for more information. ## Real Example: When Sourcing Doesn't Work **Author Note: Inflection's Experience** Inflection tried sourcing tools. We were exactly the kind of fund that thought we should: thesis-driven, research-focused, proactive outreach. But the companies we invested in were too early to exist in datasets (pre-seed deep tech, often not even incorporated yet), and our criteria were too specific (not just "deep tech" but particular themes). Broad searches created overwhelming noise. Our best deal flow came from research that attracted inbound and relationships that led to warm intros. ## The Bottom Line Sourcing is where every fund wants to start, but question whether you actually need it first. Many early-stage, thesis-driven funds find that companies are too early for datasets and criteria are too specific for effective filtering. For funds where sourcing works (Series A+, broad criteria), the tools are commoditized, don't build. The only exception is genuinely unique niches like relationship graph optimization or specialized signal detection. Don't start with sourcing just because everyone else does. In the next chapter, we'll look at CRM and deal flow management, which is more universally valuable across fund types. # Website and External Presence Source: https://buildingfor.vc/guide/part-2-tech-stack/website-and-external-presence How to approach your fund's website and broader online presence without overcomplicating the technology. ## Overview Your website and external presence are how founders and LPs discover you. Before a founder takes your call, they look at your website. Before an LP considers your fund, they research your online presence. These are first impressions that matter enormously. This is less about technology and more about branding, design, and content. The technology should be simple and stay out of the way. What matters is that your website clearly communicates what your fund is about, looks professional, and makes it easy for people to understand if you're the right investor for them. Your external presence extends beyond your website to published research, social media, and anywhere else people encounter your fund online. This ties directly back to [research platforms](/guide/part-2-tech-stack/research-platforms) - publishing your thinking is how you demonstrate expertise and attract the right founders. This chapter covers how to approach your website, what technology choices make sense, and how to think about your broader external presence without overcomplicating it. ## Your Website Your fund's website is often the first thing founders or potential LPs see. It's worth investing in getting it right. **Hire design help**: Unless you have design expertise on your team, hire a brand agency to design your website. The difference between a professionally designed site and one you design yourself is immediately obvious. Founders and LPs are forming judgments based on these first impressions. At Inflection, we worked with an external brand agency to design the website. They handled visual design, branding, and user experience. We implemented it ourselves using their designs. The results were much better than if we'd tried to design it internally. Good design communicates professionalism and attention to detail before anyone reads a word of your content. **Use a modern CMS**: Choose a content management system that makes it easy to update your website without editing code. You'll get frequent requests to update content: adding new investments, updating team member bios, changing copy, publishing new content. If updating the website requires a developer every time, it becomes a bottleneck. At Inflection, we used [Sanity](https://www.sanity.io/). It's a modern headless CMS that lets team members update content through a clean interface. When we made a new investment, anyone on the team could add it to the portfolio page. When we wanted to update copy, it was a quick edit in Sanity rather than a code change. This is the kind of practical convenience that matters day-to-day. Other good options include [Contentful](https://www.contentful.com/) (another headless CMS) or even [WordPress](https://wordpress.com/) if you want something established and familiar. The key is that non-technical team members should be able to update content themselves. **What your website should communicate**: Your fund's investment focus. What stages do you invest at? What sectors or technologies? What geography? Founders should know immediately if you're relevant to them. Your portfolio. Show your investments prominently. This builds credibility and helps founders understand what types of companies you back. Your team. Who makes investment decisions? What are their backgrounds? LPs and founders both care about this. Make it easy to understand who they'd be working with. Your thinking. If you publish research or have strong points of view, feature this content. It demonstrates expertise and helps attract founders in spaces you understand deeply. How to reach you. Make it easy for founders to get in touch. Clear contact information, ideally a simple way to submit a pitch or introduction request. **Branding is essential**: Your website should reflect the type of fund you are. A deep tech fund focused on technical founders should feel different from a consumer fund focused on brand and community. Your visual identity, tone, and content choices all communicate who you are. This is why design help matters - good designers understand how to make these choices cohesively. ## Publishing Research and Thought Leadership This ties directly to [research platforms](/guide/part-2-tech-stack/research-platforms). If you're doing deep research on markets and technologies, publishing that research builds your brand and attracts founders. **Why publish**: Publishing demonstrates expertise in your focus areas. When a founder is building in infrastructure software and they read your detailed analysis of the infrastructure market, they're more likely to reach out to you. Publishing is how you show founders you understand their space before they ever talk to you. Publishing also attracts inbound deal flow. Founders building in spaces you've researched see your content and realize you're a relevant investor. This is proactive deal generation through content rather than through systematic sourcing. **Where to publish**: You have several options. Many funds use [Substack](https://substack.com/) for regular content. It's simple, familiar to readers, and lets you build an email list of subscribers. Others publish on [Medium](https://medium.com/), which has built-in distribution through its platform. Some funds publish directly on their website using their CMS. The platform matters less than the consistency and quality of the content. If you're publishing thoughtful research regularly, the specific platform doesn't make much difference. **What to publish**: The research you're already doing. Market analysis, technology trends, thesis pieces about where you're investing and why. Turn your internal research into external content. If you're researching AI infrastructure deeply, publish that research. It demonstrates your thinking and attracts companies in that space. The key is that publishing should come from real research, not be forced marketing content. Founders can tell the difference. If your published content reflects genuine deep thinking about markets and technologies, it works. If it's shallow content created just to have something to publish, it doesn't help. ## Social Media and Other Channels Beyond your website and published research, social media is where much of your external presence lives. **LinkedIn and X**: These are table stakes for any fund. Share portfolio news, team updates, published research, and relevant industry content. Like any company, you want to be active and engaged on these platforms. Individual partners having their own presence on LinkedIn and X often matters more than the fund's official accounts. When partners share their thinking, make introductions, and engage with their networks, it builds the fund's reputation through their personal brands. **Other platforms**: Some funds use Instagram, TikTok, or other platforms. This usually makes sense only if you have dedicated marketing resources. Instagram works well if you're building a brand around community or visual content. But it requires consistent effort and content creation. Unless you have a marketing person focused on this, don't feel pressured to be on every platform. Focus on LinkedIn and X where your target audience (founders and LPs) already are. Do those well before expanding to other channels. **The technology perspective**: There's not much technology infrastructure needed for social media presence. You're using the platforms themselves. The work is creating good content and engaging consistently. Scheduling tools ([Buffer](https://buffer.com/), [Hootsuite](https://www.hootsuite.com/)) can help if you're managing multiple accounts, but the core work is content and community, not technology. ## The Bottom Line Your website and external presence are about first impressions and credibility. They're less about technology and more about branding, design, and content. Hire design help for your website. Use a modern CMS so non-technical team members can update content without developer involvement. Focus on clearly communicating your investment focus, portfolio, team, and how to reach you. Publishing research builds your brand and attracts founders in spaces you understand deeply. Use Substack, Medium, or your own blog. The platform matters less than the quality and consistency of content. Publish genuine research, not forced marketing content. For social media, focus on LinkedIn and X where founders and LPs already are. Individual partner presence often matters more than official fund accounts. Don't feel pressured to be on every platform unless you have dedicated marketing resources. The technology is simple. The work is creating content that demonstrates expertise and building a presence that attracts the right founders. In the next chapter, we'll bring everything together: how to sequence your tech stack implementation across three phases. # Choosing Your Technology Stack Source: https://buildingfor.vc/guide/part-3-technical-foundations/choosing-your-stack The default technology stack for VC tools, when to deviate from it, and why shipping fast matters more than using the newest framework. ## Overview Your job is to ship tools that help your fund make better investment decisions. You're building internal tools for 5-20 people, not the next Stripe. The right stack lets you ship fast, maintain easily, and integrate with VC tools like PitchBook and Attio. This chapter covers the default technology stack for VC, when to deviate from it, and why velocity matters more than architecture. ## Why Technology Choices Matter for VC The technology decisions you make have direct consequences for what you can accomplish: **Velocity matters more than performance**. You're building tools for tens of users, not millions. A research platform that loads in 200ms versus 50ms makes no difference to a GP. But a research platform that exists in 2 weeks versus 2 months makes a huge difference to whether you get support for the next project. **Maintenance matters more than elegance**. You'll be maintaining this code, probably alone, for years. The clever architecture that seemed smart when you built it becomes a nightmare when you're the only person who understands it. Boring, well-documented, widely-used technology means you can fix things quickly and onboard others if your team grows. **Integration matters more than custom solutions**. VC funds live in an ecosystem of specific tools: Attio or Affinity for CRM, PitchBook for data, Harmonic or Specter for sourcing. Choosing technology that makes API integrations easy is more valuable than choosing technology that's theoretically more powerful. **AI coding tools work better with popular frameworks**. [Claude Code](https://code.claude.com/docs/en/overview), [Cursor](https://cursor.com/), and similar tools dramatically accelerate development with well-documented stacks. This makes choosing boring, widely-used technology even more valuable. **Your time is the constraint, not compute**. Infrastructure costs for internal VC tools are trivial compared to data subscriptions and your salary. Optimize for developer velocity, not hosting costs. ## The Default Stack At Inflection, our default stack for web applications and APIs was: * **[Next.js](https://nextjs.org/)**: React framework with built-in routing, API routes, server-side rendering * **[TypeScript](https://www.typescriptlang.org/)**: JavaScript with static typing * **[shadcn/ui](https://ui.shadcn.com/)**: Component library built on Radix UI and Tailwind CSS * **[PostgreSQL](https://www.postgresql.org/)**: Relational database * **[Neon](https://neon.com/)**: Serverless Postgres * **[Drizzle](https://orm.drizzle.team/)**: TypeScript ORM for Postgres * **[Vercel AI SDK](https://ai-sdk.dev/)**: For LLM integrations * **[Deployed on Vercel](https://vercel.com/)**: Git push equals deployment, zero DevOps overhead This isn't the only viable stack. But it's the right default for most VC tools, and here's why in VC-specific terms: **Next.js gives you full-stack in one framework**. You can build UI, API endpoints, background jobs, and webhooks all in the same codebase. For a research platform or sourcing tool, this means you don't need separate frontend and backend repos, separate deployments, or coordination between different frameworks. One codebase, one deployment, less complexity. **TypeScript catches errors before they reach users**. When you're shipping fast with limited testing, static typing is your safety net. The compiler catches bugs that would otherwise break production. And with modern AI coding tools, TypeScript's explicitness helps Claude Code understand your codebase and generate correct code. **shadcn/ui provides good-enough UI components**. VCs aren't evaluating your tools based on design polish. They care about functionality and insights. Shadcn gives you professional-looking components that work well without hiring a designer. Your GPs will never complain that you used shadcn/ui instead of a custom design system. **Postgres is the right database for most VC tools**. It's relational (your data is structured: companies, deals, people, relationships), it's mature, it scales far beyond what you'll need, and as covered in [Data Warehousing](/guide/part-3-technical-foundations/data-warehousing), it's adequate as a data warehouse for small to mid-size funds before you need Snowflake or BigQuery. **Vercel deployment means zero DevOps**. You push to git, Vercel builds and deploys, you get a URL. No Docker, no Kubernetes, no CI/CD configuration, no server management. For internal tools where you're the only developer, eliminating DevOps overhead is essential. **AI SDK makes LLM integration trivial**. Most tools you build will use LLMs for something: summarizing companies, analyzing research, answering questions. Vercel's AI SDK handles streaming, function calling, and failover between providers. You focus on prompts and features, not plumbing. **AI coding agents work best with this stack**. Next.js and TypeScript are among the most popular technologies developers use. This means Claude Code, Cursor, and Copilot have seen enormous amounts of training data with these frameworks. They generate better code, make fewer mistakes, and help you ship faster than if you chose something obscure. **Adding Skills to your agent**: You can extend Claude Code or Codex with Skills that teach them best practices for specific frameworks. Browse [skills.sh](https://skills.sh/) for inspiration. For example, [Vercel React Best Practices](https://skills.sh/vercel-labs/agent-skills/vercel-react-best-practices) helps your agent write performant React code, [Web Design Guidelines](https://skills.sh/vercel-labs/agent-skills/web-design-guidelines) improves UI decisions, and [Supabase Postgres Best Practices](https://skills.sh/supabase/agent-skills/supabase-postgres-best-practices) guides database schema design. These help your agent make better decisions as it writes code for your app. This stack isn't exciting. It's boring, mature, widely-used technology. That's exactly what you want. Your value comes from what you build, not from using cutting-edge frameworks. ## When to Introduce New Technology The default stack covers most of what you'll build. But sometimes you need something it doesn't provide. The rule is: pick one or two new technologies per project, not more. **For Kepler** (Inflection's research platform), we introduced **[Tiptap](https://tiptap.dev/)** to create a Notion-like editing experience. This made sense because: * The editing experience was core to the product's value * Tiptap is mature, well-funded, and widely used * It integrated cleanly with React/Next.js * It solved a problem the default stack didn't address We didn't also introduce a new database, a new deployment platform, a new state management library, and a new UI framework. We changed one thing where it mattered. **Evaluate new technology with these questions:** 1. **Does the default stack actually not solve this?** Often you can build what you need with what you already have. Don't add complexity for marginal improvements. 2. **Is this technology mature and well-maintained?** Check recent commits, number of contributors, whether there's a company behind it or just a solo maintainer. You don't want to adopt something that will be abandoned in 6 months. 3. **Does it have financial backing?** If it's open source, is there a company funding development? If it's proprietary, is the company sustainable? Switching technologies later is expensive. 4. **Does it integrate well with your existing stack?** Adding a Python-based tool to an otherwise TypeScript codebase creates friction. Not insurmountable, but adds complexity. 5. **Will AI coding agents understand it?** Obscure libraries mean Claude Code can't help you as effectively. Popular libraries mean you ship faster. For most projects, you won't introduce new technology. When you do, be selective and strategic about it. ## When Python Makes Sense The default stack for us is TypeScript, but you can't escape Python entirely at a VC fund. Here's where Python is the right choice: **Data analysis and exploration**. When a GP asks "how many Series A companies in fintech raised in the last 6 months?", you open a [Jupyter](https://jupyter.org/) or [Hex](https://hex.tech/) notebook, write SQL to pull data from your warehouse, and use [pandas](https://pandas.pydata.org/) to analyze it. Python notebooks are the standard for ad-hoc data work. TypeScript isn't competitive here. **Data transformations in the warehouse**. If you're using [dbt](https://www.getdbt.com/) for data transformations (covered in [Data Warehousing](/guide/part-3-technical-foundations/data-warehousing)), you'll write Python for custom models that run in [Snowpark (Snowflake)](https://www.snowflake.com/en/product/features/snowpark/), [BigQuery DataFrames](https://docs.cloud.google.com/bigquery/docs/use-bigquery-dataframes), or [PySpark](https://spark.apache.org/docs/latest/api/python/index.html). The data warehouse ecosystem is Python-native. **Machine learning and AI work**. If you're building anything involving model training, embeddings, or scientific computing, Python's ecosystem ([scikit-learn](https://scikit-learn.org/stable/), [numpy](https://numpy.org/), [PyTorch](https://pytorch.org/)) is unmatched. **Our approach**: TypeScript by default for everything. Python when you genuinely need it for data work. That's two stacks total: web stack (TypeScript/Next.js) and data stack (Python). Don't add a third unless you have an exceptional reason. Interestingly, TypeScript has gotten much better for data purposes, but you still can't escape Python for notebooks, dbt, and in-warehouse computation. ## Existing Code: Don't Rewrite Unless It's Broken A common scenario: you join a VC fund that already has tools built by someone who left. The code is in a different stack than you'd choose. It's not documented well. It's not how you'd architect it. Your first instinct is to rewrite it in your preferred stack. **Resist this instinct.** If the existing tools work and serve GPs well, continue building on them. Your first job isn't to rip everything out and start fresh. It's to understand what the fund needs, fix what's broken, and add value incrementally. **Only rewrite existing code if:** * **It's blocking new features**. The architecture is so inflexible that you can't add what GPs need without a rewrite. * **It's unmaintainable**. The code is so poorly structured or undocumented that fixing bugs takes weeks instead of hours. * **It's actively broken**. The tools don't work, GPs don't use them, and fixing them would take longer than rebuilding. **If you do rewrite, do it incrementally.** Don't announce "I'm rewriting everything, it'll be ready in 3 months." Build the new version alongside the old one, migrate features one by one, and deprecate the old system only when the new one is proven. This keeps GPs working while you improve infrastructure. At Inflection, I was blessed to have no legacy code, but if we'd inherited working tools, the default would be to keep them running and iterate rather than rebuild. ## Lean on Managed Services Use managed services for everything unless you have a specific reason not to. **For hosting**: [Vercel](https://vercel.com/) for Next.js apps. [Railway](https://railway.com/), [Render](https://render.com/), or [Fly.io](https://fly.io/) for other services. These platforms handle deployment, scaling, monitoring, and uptime. You push code, they run it. The alternative is configuring AWS infrastructure, managing Docker containers, setting up CI/CD, monitoring servers, and handling incidents. Not worth your time. **For databases**: [Neon](https://neon.com/), [Supabase](https://supabase.com/), or [PlanetScale](https://planetscale.com/) for Postgres. [AWS RDS](https://aws.amazon.com/rds) if you're already committed to AWS. Don't run your own database servers. Let someone else handle backups, replication, and uptime. **For background jobs**: If you need async task processing, use a managed service like [Trigger.dev](https://trigger.dev/), [Inngest](https://www.inngest.com/), [Vercel's workflow SDK](https://vercel.com/docs/workflow) rather than running your own job queue infrastructure. **For monitoring**: Use Vercel's built-in monitoring or services like [Sentry](https://sentry.io/) for error tracking. Don't build your own observability stack. **The cost argument**: Managed services are more expensive than self-hosting. Vercel costs more than running an EC2 instance. Managed Postgres costs more than running your own RDS instance. But the cost difference is trivial compared to your time. Your engineering time, at a fully-loaded cost including salary and benefits, is worth (at minimum) \$100-200/hour at a VC fund. Spending 4 hours a month managing infrastructure to save \$200 on hosting is a terrible trade. Spend that time building features that help GPs make better investment decisions. **The exception**: If you have specific compliance requirements ([Security and Compliance](/guide/part-3-technical-foundations/security-and-compliance) covers security and compliance), you might need self-hosted infrastructure for data residency or audit reasons. But this is rare for most funds and work with your CISO or compliance officer to find a path forward. Default to managed services. ## Dependencies: Be Selective But Not Afraid Minimize dependencies, but don't be dogmatic about it. If a well-maintained library solves your problem and lets you ship faster, use it. **In the VC ecosystem, you'll use vendor SDKs and integrations:** * API wrappers from your data providers * Vercel AI SDK for LLM integrations * Drizzle or Prisma for database access These are core dependencies that make sense. They're maintained by companies who have financial incentive to keep them working, and they save you from writing API integration code from scratch. **Be selective about other dependencies:** * A well-maintained date library (date-fns) saves time and reduces bugs * A validation library ([Zod](https://zod.dev/) for TypeScript, covered in [Integrations and APIs](/guide/part-3-technical-foundations/integrations-and-apis)) is essential for API integrations * A testing library ([Vitest](https://vitest.dev/)) makes sense if you write tests (though for internal tools, extensive testing is often overkill) **Be wary of:** * Dependencies that do things you could write in 10 lines of code (micro-dependencies) * Dependencies with no recent commits or single maintainers * Dependencies that pull in huge transitive dependency trees Check financial backing and maintenance status. A library with a company behind it is safer than a side project from a solo developer. Not a hard rule, but a useful signal. ## What Actually Matters GPs don't care that you used Next.js or have elegant architecture. They care about whether your tools help them find better companies, make better decisions, and generate better returns. Your value comes from building research platforms that surface insights, sourcing tools that identify companies early, and analyses that inform thesis development. Not from using the newest framework or optimizing performance beyond what's necessary. **Modern AI tools change everything**. With Claude Code or Cursor, learning any language or framework is fast. AI agents handle syntax, suggest patterns, and catch mistakes. This means two things: **For you**: Choose popular, well-documented frameworks. The AI tools work dramatically better with them, letting you ship faster. **For hiring**: Prioritize VC domain knowledge and product sense over specialized technical expertise. For more on hiring, see [Hiring Your Data Team](/guide/part-1-understanding-vc/hiring-your-data-team). Technology is a means to an end. The end is helping your fund deploy capital more effectively. Choose technology that lets you reach that end faster. ## The Bottom Line Use boring, proven technology that lets you ship fast. At Inflection, that was Next.js, TypeScript, Postgres, and Vercel. For data work, Python when you need it. Only introduce new technology (like Tiptap for Kepler) when it solves a problem the default stack doesn't address. Pick one or two new technologies per project, not more. Ensure they're well-maintained and financially backed. If you inherit existing code that works, continue building on it. Don't rewrite unless it's truly blocking progress or broken beyond repair. Lean on managed services for everything. Your time is worth more than the cost difference between managed and self-hosted infrastructure. Be selective about dependencies, but use well-maintained libraries that speed up development. In VC, that includes vendor SDKs for CRMs, data providers, and LLM integrations. With modern AI coding tools, learning new languages and frameworks is fast. Language expertise matters less than VC domain knowledge and product sense. Most importantly: GPs care about insights and decisions, not technology. Your value comes from what you build, not how you build it. Ship fast, prove value, get support for bigger projects. Use technology that enables that, not technology that impresses other engineers. In the next chapters, we'll cover data providers: the external vendors who supply information about companies, funding, people, and markets that you'll integrate into the tools you build. # Data Modeling for Venture Capital Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-modeling How to structure VC data, the critical distinction between companies and deals, and practical patterns for modeling the investment lifecycle. ## Overview How you model your data determines what you can analyze, how fast you can build features, and whether your systems adapt as needs evolve. This chapter covers core entities, the critical distinction between companies and deals, temporal data handling, and practical patterns from building real VC systems. ## Core Entities and Relationships Your data model will vary based on what you're building (CRM, research platform, portfolio dashboard), but most VC systems share common entities. **Companies** - Startups you might invest in, portfolio companies, competitors. Core attributes: name, description, website, founding date, location, sector, stage. Companies are the center of your data model. **Deals** - Your fund's relationship with a company (covered in detail in the next section). One company can have multiple deals: you might pass on their seed round, then invest in their Series A. **People** - Founders, executives, employees. Core attributes: name, email, LinkedIn URL. People connect to companies through roles, not direct relationships. **Roles** - The connection between people and companies capturing who works where, in what capacity, and when. Model as a separate entity to track full history: a founder worked at Google (2018-2020), started Company X (2020-present), while advising Company Y (2021-present). **Education** - Where people went to school, what they studied, when. Model separately to capture multiple degrees and overlapping education. Enables analysis like "which universities produce the most founders in our focus areas?" **Funding Rounds** - Capital raised by companies. Attributes: round type, amount, valuation, dates, investors. Expect messiness: unreported SAFE notes, incomplete data from vendors, unannounced rounds. Your data model needs to handle uncertainty (covered in "Dealing with Messy Reality" below). **Investors and Funds** Other VC firms, angels, corporate investors who participate in rounds. You care about investors for co-investment patterns, warm intro paths, and understanding who's active in your sectors. **Critical: Model the hierarchy as Investor → Fund → Funding Round**, not just Investor → Funding Round. Here's why this matters: One investor (the firm) can have multiple funds. Sequoia Capital has multiple funds. a16z has multiple funds. Your own firm probably has multiple funds (Fund I, Fund II, etc.). When tracking who invested in a company, you need to know which specific fund made the investment, not just which firm. This matters because different funds from the same investor can make different investment decisions. Inflection Mercury Fund might pass on a deal while Inflection Mars Fund invests. If you model this as just "Inflection invested," you lose critical information. The SQL below is simplified to illustrate entity relationships - not production-ready code. ```sql theme={null} -- Investors table (the firms) CREATE TABLE investors ( id UUID PRIMARY KEY, name TEXT NOT NULL, -- "Sequoia Capital" type TEXT, -- VC, Angel, Corporate headquarters TEXT, focus_areas TEXT[] ); -- Funds table (specific funds within firms) CREATE TABLE funds ( id UUID PRIMARY KEY, investor_id UUID REFERENCES investors(id), name TEXT NOT NULL, -- "Sequoia Capital Fund XIV" vintage_year INTEGER, fund_size NUMERIC, status TEXT -- Active, Deployed, Closed ); -- Funding round participations (which funds invested in which rounds) CREATE TABLE funding_round_participants ( id UUID PRIMARY KEY, funding_round_id UUID REFERENCES funding_rounds(id), fund_id UUID REFERENCES funds(id), -- Links to specific fund, not just investor amount_invested NUMERIC, is_lead BOOLEAN ); ``` Some data providers (some are better at this than others) provide fund-level detail. When available, capture it. When it's not available (data provider only knows "Sequoia invested" but not which Sequoia fund), you can create a placeholder fund or link directly to the investor, but structure your schema to support fund-level data when you have it. This distinction has bitten many VC data systems. Get it right from the start. **LPs (Limited Partners)** - Individuals and institutions who invest in your fund. Required for fund operations tools or LP portals. LPs can be individuals (HNWIs with personal info) or institutions (pension funds, endowments with organizational info). Model with a type field and nullable columns, or separate tables if managed differently. The list above isn't exhaustive. Depending on what you're building, you might also model: board seats, advisors, customers (for B2B portfolio companies), competitors, acquisitions, exits, fund entities (if you manage multiple funds), and more. Start with the entities your application actually needs, not a comprehensive schema you think you might need someday. ## Companies vs. Deals: The Critical Distinction The most important modeling decision in VC systems is separating companies from deals. **A company is an entity that exists independently of your fund.** Stripe is a company. It has founders, funding rounds, employees, a product, customers. Stripe exists whether or not your fund ever talks to them. **A deal is your fund's relationship with a company.** You sourced Stripe, you met with the founders, you decided to invest (or pass), you negotiated terms, you closed the investment. The deal represents your fund's interaction with that company. **Why separate them?** Your fund can have multiple deals with the same company. You might: * Pass on their seed round (Deal 1: Status = Passed) * Invest in their Series A two years later (Deal 2: Status = Closed) * Consider a follow-on in their Series B (Deal 3: Status = In Diligence) Each deal has different properties: deal source, who at your fund is leading it, what stage it's at, when you first talked to them, what your investment thesis was, what terms you negotiated. These properties belong to the deal, not the company. Meanwhile, company properties (employee count, funding history, product description) are shared across all deals. If you update Stripe's employee count, that change applies regardless of which deal you're looking at. **In practice, this means separate tables:** ```sql theme={null} -- Companies table CREATE TABLE companies ( id UUID PRIMARY KEY, name TEXT NOT NULL, website TEXT, description TEXT, founded_date DATE, location TEXT, sector TEXT, stage TEXT, created_at TIMESTAMP, updated_at TIMESTAMP ); -- Deals table CREATE TABLE deals ( id UUID PRIMARY KEY, company_id UUID REFERENCES companies(id), -- Foreign key to company deal_name TEXT, -- "Stripe Series A" status TEXT, -- Sourced, Meeting, Diligence, IC, Term Sheet, Closed, Passed source TEXT, -- How we found this company (warm intro, Harmonic, conference) lead_investor UUID REFERENCES people(id), -- Who at our fund is leading initial_contact_date DATE, last_activity_date DATE, investment_amount NUMERIC, ownership_percentage NUMERIC, created_at TIMESTAMP, updated_at TIMESTAMP ); ``` When you query deals, you join to companies to get company information. When you're analyzing companies (building market maps, tracking sectors), you query companies directly without caring about deals. This separation is fundamental to building VC software that works. If you try to embed deals into the companies table or vice versa, you'll create a mess that's hard to query and impossible to extend. As emphasized in [CRM and Deal Flow](/guide/part-2-tech-stack/crm-and-deal-flow) on CRM, this distinction is one of the most important architectural decisions you'll make. ## Dealing with Messy Reality VC data is messy. Your data model needs to account for this rather than assuming perfect information. **Don't assume GPs make one investment per company** As covered above, model deals separately from companies. Your fund might invest multiple times, or pass then later invest, or invest then later decline a follow-on. The data model needs to support multiple deals per company. **Don't assume perfect funding round data** Early-stage companies raise money in ways that aren't always captured in data sources. SAFE notes, convertible notes, rolling closes, stealth rounds. Crunchbase and PitchBook miss things, especially for pre-seed and seed companies. Your data model should allow for: * Unknown or estimated funding amounts * Approximate dates (Q2 2024, not a specific day) * Rounds that might not exist (rumored but unconfirmed) * Multiple rounds of the same type (Seed, Seed Extension, Seed II) Don't build validation rules that enforce "Series B must come after Series A" or "funding amounts must be known." Real companies break these rules constantly and your data sources, even more so. **People have overlapping experiences and education** Founders often work multiple jobs simultaneously: running their startup while advising other companies, sitting on boards, teaching part-time. They pursue education while working: part-time MBAs, executive programs, online courses. This is why roles and education must be separate tables with start and end dates, not single fields in the people table. You need to capture overlapping time periods. ```sql theme={null} -- Roles table (separate from people and companies) CREATE TABLE roles ( id UUID PRIMARY KEY, person_id UUID REFERENCES people(id), company_id UUID REFERENCES companies(id), title TEXT, role_type TEXT, -- Employee, Founder, Advisor, Board Member start_date DATE, end_date DATE, -- NULL if current is_current BOOLEAN, created_at TIMESTAMP ); -- Education table (separate from people) CREATE TABLE education ( id UUID PRIMARY KEY, person_id UUID REFERENCES people(id), institution TEXT, degree TEXT, -- BS, MS, MBA, PhD field_of_study TEXT, start_date DATE, end_date DATE, created_at TIMESTAMP ); ``` With this structure, you can query "show me all founders who worked at Google" (join people → roles → companies where company = Google and role\_type = Employee). You can analyze "which schools produce the most founders in fintech" (join people → roles → companies → education, filter for fintech sector and founder roles). If you tried to store current role and education as fields in the people table, you'd lose history and couldn't answer these questions. ## Temporal Data and History Companies change constantly. Employee count grows. Funding rounds happen. Valuations change. Products pivot. You need to decide what history to keep and how to model it. **Start with append-only tables** The simplest approach: never update or delete data. When you get new information about a company, append a new row with a timestamp. This creates an audit trail of every change. ```sql theme={null} -- Append-only company data from PitchBook CREATE TABLE pitchbook_company_changes ( id UUID PRIMARY KEY, company_id UUID, fetched_at TIMESTAMP, -- When we got this data raw_json JSONB, -- Raw response from PitchBook API name TEXT, description TEXT, employee_count INTEGER, funding_total NUMERIC, -- ... other fields created_at TIMESTAMP ); ``` Every time you fetch company data from PitchBook (daily, weekly, or on-demand), you insert a new row. You never update existing rows. This preserves full history: you can see what PitchBook reported about a company on any date. **Use dbt to transform to latest values** Append-only tables are great for history but inconvenient to query. You usually want the latest information, not every historical value. Use [dbt](https://www.getdbt.com/) (covered in [Data Warehousing](/guide/part-3-technical-foundations/data-warehousing)) to transform append-only staging tables into intermediate tables with current values: ```sql theme={null} -- dbt model: models/intermediate/pitchbook_companies.sql -- Gets the latest record per company from append-only staging data SELECT DISTINCT ON (company_id) company_id, name, description, employee_count, funding_total, fetched_at FROM {{ ref('pitchbook_company_changes') }} ORDER BY company_id, fetched_at DESC ``` Now you have both: `pitchbook_company_changes` (full history) and `pitchbook_companies` (latest values). Applications usually query the latest values table. When you need to analyze "how did this company change over time?" you go back to the append-only table or create a custom table for those types of queries. **What history actually matters** Storage is cheap these days. Rather than deciding what history to keep and what to discard, it's often simpler to just keep everything in your append-only staging tables. A year of daily company snapshots for thousands of companies costs pennies in Postgres or cloud data warehouse storage. The practical considerations are query performance and data warehouse costs (if you're using Snowflake or BigQuery where you pay per query). But even there, you're usually querying the transformed and materialized "latest values" tables, not the full historical append-only tables. **Keep all history in staging tables.** When you need to analyze "how did this company change over time?" you have the data. When you don't need it, you're not querying it, so it doesn't cost anything. This is simpler than trying to decide upfront what historical data might be valuable someday. The main exception: if you're storing truly high-volume data (real-time metrics, streaming data, logs), you might need retention policies. But for standard VC data (company attributes, funding rounds, people roles), just keep it all. As you add more data sources, follow dbt's layered approach (staging → intermediate → marts) to stay organized. The [Data Warehousing](/guide/part-3-technical-foundations/data-warehousing) chapter covers this methodology in detail. ## The Bottom Line Accept that VC data is messy: multiple investments per company, imperfect funding data, overlapping experiences. Your data model should handle uncertainty and incomplete information. Separate companies from deals. Use append-only tables for history (storage is cheap), transform to latest values for queries. Start with what you need and add entities as your systems grow. In the next chapter, we'll cover more advanced entity resolution: matching the same company across different data sources when they use different names, formats, and identifiers. # Accessing Data Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-providers/accessing-data How data providers deliver data: APIs, file exports, and hybrid approaches. ## Overview Data providers deliver data in two main ways: APIs and files. Each has different implications for how you build integrations and what you pay. ## API-Based Delivery Most modern vendors provide REST APIs. You authenticate with an API token, make HTTP requests, and receive JSON responses. **Benefits:** * Query exactly what you need (don't download everything) * Easy to integrate into applications (CRM enrichment, sourcing tools) * Can build interactive features (search, live updates) **Considerations:** * Rate limits constrain how fast you can query * Per-request or per-entity costs (each API call might cost money) * Need to handle failures, retries, validation * API schemas change over time ## File-Based Delivery Almost all vendors also provide data as file exports: CSV, [Parquet](https://parquet.apache.org/), or JSON files that you download or they upload to your S3 bucket. This is common for bulk data (entire company database, historical funding data, periodic dumps). However, this is often in the "premium" tier, which is usually many times more expensive than API access. **Benefits:** * Cheaper than API calls if you need the full dataset * Get everything at once (good for analytics, data warehouse loading) * Predictable costs (usually flat fee for the subscription) * No rate limits once you have the file **Considerations:** * Data is a snapshot (might be stale compared to API data) * Need to process and load files * Need to handle incremental updates ### Prefer Parquet Over CSV Vendors will sometimes give you the option of CSV, TSV, JSON/JSONL or [Parquet](https://parquet.apache.org/). **Always choose Parquet**. It includes schema information (you know data types without guessing), compresses well (smaller files), and loads much faster into data warehouses. JSON/[JSONL](https://jsonlines.org/) is also a decent choice, but typically more cumbersome to work with. CSV/TSV files require parsing, have encoding issues, no schema, and are slower to work with. Ask vendors to provide Parquet if they don't already. Most modern vendors support it. ## Hybrid Approaches Some vendors offer both APIs (for real-time enrichment) and bulk exports (for loading your data warehouse). This is ideal: use the API for interactive features, use bulk exports to populate your data warehouse efficiently. ## File Delivery Authentication For file-based vendors, there are three common options: * They upload files to their own S3 bucket that you can access (they provide AWS credentials) * They upload files directly to your S3 bucket (you provide them with write-only credentials) * They provide download links through their portal (you download manually or via script) The third option (manual download) doesn't scale well. Prefer vendors who can integrate with your cloud storage. ## Working with APIs For details on API authentication, validation, rate limits, error handling, and other integration patterns, see [Integrations](/guide/part-3-technical-foundations/integrations-and-apis). # Company Data Providers Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-providers/company-data Providers for company information, funding history, and investment data. ## Overview Company and investment data is the foundation of VC technology. You need information about companies (name, location, founding date, description, website) and their funding history (rounds, amounts, dates, investors, valuations). This data powers your CRM, deal flow tracking, market analysis, and portfolio monitoring. ## What You're Looking For **Company basics:** * Name, description, website * Location and founding date * Industry/sector classification * Employee count **Funding history:** * Investment rounds (dates, amounts, valuations) * Investors in each round * Lead investors and co-investors * Total funding raised **Investor data:** * Which VCs invested in which companies * Fund-level information * Investment patterns and focus areas ## Providers | Provider | Strengths | Weaknesses | Price | | ----------------------------------------- | ------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | -------- | | [PitchBook](https://pitchbook.com/) | Analyst-verified financials, valuation history, cap tables, detailed deal terms. Deep PE/VC data. | Overkill for early-stage focused funds. Dashboard-first, API secondary. | \$\$\$\$ | | [Crunchbase](https://www.crunchbase.com/) | Broad startup coverage, user-friendly UI, AI-powered search, good for prospecting. | Less depth on financials and deal terms. User-contributed data can be stale/inaccurate. | \$\$\$ | | [Dealroom](https://dealroom.co/) | Strong European coverage, government partnerships, ecosystem mapping, early-stage depth. | Weaker US coverage compared to alternatives. | \$\$\$ | | [Tracxn](https://tracxn.com/) | Worldwide coverage (better than others in emerging markets), impressive sector granularity | More reliance on automated data collection, less human verification | \$\$\$ | ## Alternatives The providers above involve human curation and verification, making them more accurate and contextually rich. The alternatives below are primarily scraped from online sources: useful for specific use cases and supplementary data, but typically less depth and accuracy. | Provider | Description | | --------------------------------------------------------------- | --------------------------------------------------------------------------------------- | | [CrustData](https://crustdata.com/) | Real-time firmographics via API. Headcount, funding, web traffic signals. | | [People Data Labs](https://www.peopledatalabs.com/company-data) | People data provider that also offers company firmographics. | | [Mixrank](https://mixrank.com/datasets/company-data/) | Basic firmographics. | | [Coresignal](https://coresignal.com/company-dataset/) | Comprehensive firmographics with many data points. Traditionally known for people data. | More on these in the [People Data Providers](/guide/part-3-technical-foundations/data-providers/people-data) section. ## Considerations **Coverage varies by geography and stage:** Some providers are stronger in US tech, others in European markets, others in specific stages. Test coverage against your target market. **Data freshness matters:** How quickly does the provider capture new funding rounds? Some are faster than others. **API vs dashboard:** Some providers focus on dashboard access, others have better APIs for programmatic access. Know what you need. See [Considerations](/guide/part-3-technical-foundations/data-providers/considerations) for cost and vendor relationship guidance. # Data Provider Considerations Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-providers/considerations Cost, vendor relationships, and things to think about when working with data providers. ## Overview Data providers are foundational infrastructure, but they're also expensive and require ongoing management. This page covers cost considerations, how to work with vendors, and things to keep in mind. ## Cost Considerations Data providers are expensive. Budget for them appropriately and understand different pricing models. ### Pricing Models **Per-request or per-entity:** You pay for each API call or each entity returned. A person data provider might charge per record. This is predictable per query but can add up quickly with high usage. **Subscriptions:** Pay a fixed amount per year for unlimited (or high-limit) access. Common for established vendors. Costs vary widely depending on the vendor and what you're accessing. **Per-seat pricing:** Pay per user who has access (common for platforms with dashboards, less common for APIs). Usually, you buy the API access on top of per-seat pricing for data providers with per-seat pricing. ### Cost Management Strategies * **Start with what you actually need:** Don't subscribe to every vendor. Figure out your critical use cases and buy data for those first. * **Monitor spending:** Set up alerts when API usage or costs exceed thresholds. It's easy to accidentally rack up bills with per-request pricing. * **Cache data:** Don't request the same company data repeatedly. Cache responses (in your database or Redis) with appropriate TTLs. This saves money and respects rate limits. Some data providers have strict rules about how you can cache data. Always check their terms before implementing caching strategies. * **Use bulk exports when possible:** If you need to load lots of data into your warehouse, bulk exports are usually cheaper than making thousands of API requests. * **Negotiate:** Negotiate based on your expected usage, especially between vendors. Some vendors offer discounts for bulk purchases or long-term commitments. ### Budget Planning Data costs can easily exceed infrastructure costs (servers, databases, etc.). Factor this into your overall technology budget from the start (ideally before you start building out the team). Costs scale with fund size and data needs. ## Working with Vendors ### APIs Change Frequently Data vendors update their APIs more often than you'd expect. They rename fields, change data formats, add new attributes, deprecate old endpoints. Your integrations can break when this happens, but you can mitigate this risk by following best practices (see [Data Modeling](/guide/part-3-technical-foundations/data-modeling)). Pay attention to vendor communications. They usually announce breaking changes weeks in advance via email or their changelog. Set up notifications so you see these announcements. Budget time to update your integrations when schemas change. ### Collaborate with Newer Vendors If you're working with a newer data vendor, establish trust and provide feedback on their API design. They want customers to succeed and are often open to suggestions. If their API returns data in an awkward format, tell them. If they're missing fields you need, ask for them. If rate limits are too restrictive, negotiate. If they only provide CSV but you need Parquet, request it. This is win-win: they improve their product based on real usage, you get an API that's easier to work with. Established vendors are less flexible, but newer vendors appreciate detailed feedback. ### Support and Documentation Vendor quality varies significantly in support and documentation: * Some have excellent docs, responsive support, active communities * Some have minimal docs, slow support, no community * Some provide developer relations people who help with integration Factor this into vendor selection. If you're building critical infrastructure on a vendor's data, you need good documentation and reliable support. Don't just evaluate the data quality, evaluate whether you can actually build on top of it. ## Data Quality Not all data providers are equally accurate or comprehensive. Quality varies significantly between vendors and even within the same vendor for different data types. ### Accuracy vs Speed Some vendors prioritize accuracy (verify information before publishing, resulting in lag). Others prioritize speed (publish quickly, may have more errors). Choose based on your use case. ### Coverage Differences Vendors excel at different: * **Stages:** Some are better for early-stage, others for growth/late-stage * **Geographies:** Strong US coverage vs weak international * **Sectors:** Deep tech, bio, fintech specializations ### Test with Portfolio Companies The best way to evaluate vendor accuracy is to test against your portfolio companies, companies where you know the ground truth. Check if funding amounts match, if team data is current, if descriptions are accurate. This tells you what each vendor is actually good for and where they fall short. ### Data Freshness Vendors update at different cadences (real-time, daily, weekly, monthly). Know how fresh the data is and don't present stale data as current. ## The Bottom Line Data providers are foundational infrastructure. You need external data about companies, funding, people, and markets. Choose vendors based on what data you actually need. Don't subscribe to everything. Test accuracy using your portfolio companies. # Data Providers Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-providers/index Understanding external data vendors who supply information about companies, funding, people, and markets. ## Overview You need external data about companies, funding, people, and markets. Data providers feed your CRM, power sourcing tools, and enrich your warehouse. But they're expensive, have different strengths, and deliver data differently. This section covers the different types of external data for VC funds. ## Types of Data for VC Funds **Company and funding data**: Basic information about companies (name, location, founding date, description, website, employee count, industry) and their funding history (rounds, amounts, dates, investors, valuations). This is the foundation of your deal flow tracking, market analysis, and portfolio monitoring. **People data**: Information about founders, executives, and employees. Who founded the company, who's on the leadership team, their experience and education. This is crucial for evaluating founding teams and identifying potential hires for your portfolio companies. **Signal and sourcing data**: Early indicators that companies are growing, raising funding, or becoming interesting. This includes investor interest signals (who's looking at a company), company registry filings (new incorporations), talent movements, headcount growth, web traffic, and funding predictions. Useful for sourcing tools to identify companies before they're widely known. **Market and competitive data**: Information about sectors, industries, technology trends, and competitive landscapes. What companies operate in a space, how markets are evolving, what technologies are emerging. **Research data**: Academic papers, patents, and technical publications. Especially valuable for deep tech funds evaluating novel technology or scientific founding teams. Many providers span multiple categories. For example, [People Data Labs](https://www.peopledatalabs.com/) offers both person and company data. Signal providers often include company fundamentals. The pages below are (mostly) organized by primary focus, but don't assume a provider only does one thing. ## What's in This Section Recommended data stacks for different fund types and stages. APIs, file exports, authentication, and delivery methods. Providers for company information, funding history, and investment data. Providers for early-stage signals, sourcing, and growth indicators. Providers for founder, team, and employee information. Market data, research, patents, and specialized sources. Cost, vendor relationships, and things to think about. # Other Data Providers Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-providers/other-data Market data, research, patents, and specialized data sources. ## Overview Beyond company, signal, and people data, VC funds often need specialized data: market intelligence, academic research, patents, financial data, and sector-specific sources. This page isn't comprehensive, it's a starting point for inspiration. The specialized data landscape is vast and constantly evolving. If you've found a great resource that should be here, [suggest an edit](https://github.com/alexpatow/building-for-vc/edit/main/guide/part-3-technical-foundations/data-providers/other-data.mdx). ## Market and Financial Data | Provider | What It's For | Price | | -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------- | -------- | | [S\&P Capital IQ](https://www.spglobal.com/marketintelligence/en/solutions/sp-capital-iq-platform) | Comprehensive financial data. Comparables analysis, market sizing. | \$\$\$\$ | | [Bloomberg](https://www.bloomberg.com/professional/products/data/) | Real-time market data, news, analytics. Standard in finance. | \$\$\$\$ | | [Refinitiv (LSEG)](https://www.lseg.com/en/data-analytics) | Market data, news, regulatory filings. Bloomberg alternative. | \$\$\$ | | [Morningstar](https://www.morningstar.com/products/data) | Investment research, fund ratings, equity data. Strong on public markets. | \$\$\$ | These are primarily useful for growth-stage funds doing comparables analysis or funds that invest alongside public market activity. Most early-stage funds won't need this level of financial data. *** ## Research and Academic Data For deep tech and science-based funds, academic research is critical for evaluating technical founders and understanding technology landscapes. | Provider | What It's For | Price | | ---------------------------------------------------- | -------------------------------------------------------------------------------------------------- | ----- | | [arXiv](https://arxiv.org/) | Preprints in physics, math, CS, AI/ML. Track cutting-edge research before publication. | Free | | [Semantic Scholar](https://www.semanticscholar.org/) | AI-powered research discovery. Citation networks and research impact. | Free | | [Google Scholar](https://scholar.google.com/) | Broad academic search. Good for quick lookups and citation counts. Can use SerpAPI to scrape data. | \$ | | [PubMed](https://pubmed.ncbi.nlm.nih.gov/) | Biomedical and life sciences literature. Essential for bio/healthcare funds. | Free | These are mostly free and publicly accessible. The challenge isn't cost, it's knowing how to use them and having the domain expertise to interpret what you find. *** ## Patent and IP Data Patents indicate technology development, potential IP moats, and founder technical depth. | Provider | What It's For | Price | | --------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | ------- | | [Lens](https://www.lens.org/) | Patent and scholarly search, inexpensive yearly subscription for commercial use. Links patents to academic research. | \$ | | [Google Patents](https://patents.google.com/) | Quick patent searches. Good for initial lookups. Can use SerpAPI to scrape data. | \$ | | [USPTO](https://www.uspto.gov/) | Official US patent database. | Free | | [Espacenet](https://worldwide.espacenet.com/) | European patent database. Good international coverage. | Free | | [PatSnap](https://www.patsnap.com/) | Patent analytics platform. Visualization, competitive analysis. | Unknown | Most patent data is publicly available through government databases. Paid tools like Lens and PatSnap add analytics, visualization, and search on top. *** ## Web Traffic and E-commerce Data For evaluating consumer-facing companies, web traffic and e-commerce data can provide useful signals about traction and market position. | Provider | What It's For | Price | | -------------------------------------------- | ----------------------------------------------------------------------------------- | ------ | | [SimilarWeb](https://www.similarweb.com/) | Web traffic estimates, SEO data, competitive analysis. Good for consumer companies. | \$\$\$ | | [Jungle Scout](https://www.junglescout.com/) | Amazon product research, seller data, market trends. Essential for e-commerce. | \$\$ | | [Sensor Tower](https://sensortower.com/) | Mobile app analytics, downloads, revenue estimates. Good for app-based companies. | \$\$\$ | | [data.ai](https://www.data.ai/) | Mobile market data, app intelligence. Broader than Sensor Tower. | \$\$\$ | These tools are particularly useful when evaluating companies in consumer, e-commerce, or mobile-first categories where traditional funding data doesn't capture traction. *** ## Geospatial Data Location intelligence can be valuable for evaluating companies in retail, real estate, logistics, and other location-dependent sectors. | Provider | What It's For | Price | | ----------------------------------- | ------------------------------------------------------------------------------------ | ------ | | [CARTO](https://carto.com/) | Location intelligence platform. Demographics, foot traffic, site selection. | \$\$\$ | | [SafeGraph](https://safegraph.com/) | Points of interest, foot traffic patterns. Good for retail and real estate analysis. | \$\$\$ | Useful when evaluating companies where physical location matters: retail chains, logistics, real estate tech, or any business with brick-and-mortar components. *** ## Sector-Specific Data Some sectors have specialized data needs that general providers don't cover well. **Healthcare/Bio:** | Provider | What It's For | | ------------------------------------------------------------------------------------------- | --------------------------------------------------- | | [ClinicalTrials.gov](https://clinicaltrials.gov/) | Clinical trial registry. Track drug development. | | [FDA databases](https://www.fda.gov/drugs/drug-approvals-and-databases/drugsfda-data-files) | Drug approvals, safety data. | | [BioMedTracker](https://www.biomedtracker.com/) | Drug pipeline intelligence. Probability of success. | **Fintech:** | Provider | What It's For | | -------------------------------------------------- | --------------------------------------------------- | | [FDIC](https://www.fdic.gov/resources/data-tools/) | Bank data, regulatory filings. | | [SEC EDGAR](https://www.sec.gov/edgar) | Public company filings. Essential for any analysis. | **Climate/Energy:** | Provider | What It's For | | ---------------------------------------------------------- | ------------------------------------------------ | | [EIA](https://www.eia.gov/) | US energy data. Production, consumption, prices. | | [EPA databases](https://www.epa.gov/enviro/data-downloads) | Emissions, environmental compliance. | *** ## LLM-Powered Research Tools LLMs with search capabilities are increasingly useful for market research, competitive analysis, and due diligence. Unlike traditional databases, these tools synthesize information from multiple sources and return cited answers. | Provider | What It's For | Price | | -------------------------------------------- | ----------------------------------------------------------------------------------------- | ----- | | [Perplexity API](https://docs.perplexity.ai) | Search + LLM that returns sourced answers. Great for market research and quick diligence. | \$ | | [Exa](https://exa.ai/) | AI-powered semantic search. Find similar companies, research markets, discover content. | \$ | These tools are useful for: * Quick market sizing and landscape overviews * Finding competitors and similar companies * Background research on founders or technologies * Synthesizing public information during due diligence The key advantage is that they return sources with their answers, so you can verify the underlying data. Integrate them into your research workflows via API. *** ## Considerations **Specialized data requires domain expertise:** Market data and research data are only valuable if you can interpret them. These sources work best when you have domain expertise on your team. See [Considerations](/guide/part-3-technical-foundations/data-providers/considerations) for cost and vendor relationship guidance. # People Data Providers Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-providers/people-data Providers for founder, team, and employee information. ## Overview People data is critical for evaluating teams. You need information about founders (background, previous experience, education), executives (who's leading the company), and employees (team size, hiring patterns, talent quality). ## What You're Looking For **Founder and executive profiles:** * Work history and previous companies * Education and credentials * Previous exits and successes * Board positions and advisors **Team composition:** * Employee count over time * Department breakdown (engineering, sales, etc.) * Leadership team changes * Key hires and departures **Talent signals:** * Where people are moving from/to * Hiring from strong companies * Founder quality indicators ## Providers Coverage across these providers is relatively similar: not as large a gap as with company data providers. Choose based on which data points matter to you, how timely updates are, and quality of support. | Provider | Strengths | Weaknesses | Price | | --------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ | ----- | | [People Data Labs](https://www.peopledatalabs.com/) | Support has been great (personal experience). Free "cleaner" APIs have been useful in data pipelines. | Uses Elasticsearch SQL which can make filtering difficult via API. | \$\$ | | [Coresignal](https://coresignal.com/) | Recently improved, "cleaned" datasets. | | \$\$ | | [Mixrank](https://mixrank.com/) | Offers managed SQL databases for easier integration. | API can be difficult to work with. | \$\$ | ## Considerations **Data privacy:** People data has privacy implications. Understand what data you're getting and how it was collected. Some regions and websites have strict rules about processing personal data. **Accuracy varies:** Professional profiles may be outdated. Current employer might be stale. Test data quality by comparing with your portfolio or verifying manually. **Enrichment use cases:** People data is often used to enrich signal data (add founder backgrounds to company records, building a holistic view of a company's leadership team, etc.) rather than as a primary source. See [Considerations](/guide/part-3-technical-foundations/data-providers/considerations) for cost and vendor relationship guidance. # Signal Data Providers Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-providers/signal-data Providers for early-stage signals, sourcing, and growth indicators. ## Overview Signal data helps you find companies before they're widely known. These are early indicators that a company is growing, raising funding, or becoming interesting: job postings, web traffic, social media activity, product launches, hiring velocity. ## What You're Looking For **Founder signals:** * New company formations * Founders in "stealth mode" * Founder backgrounds and previous exits * Team composition changes **Funding signals:** * Indicators that a company is raising (even before announced) * Investor interest signals * Press and media coverage **Hiring signals:** * New job postings (especially leadership roles) * Hiring velocity (how fast they're growing) * Engineering team growth * Location expansion **Web and product signals:** * Website traffic growth * Product launches and updates * App store rankings * Social media activity ## Early Stage Signal Providers For pre-seed and seed investors looking for companies before they raise or become widely known. These providers often scrape data from government regulatory filings and match them with people data to round out the profile of the company. | Provider | Notes | | -------------------------------------- | -------------- | | [Gravity](https://www.gravity.inc/) | US-focused | | [Evertrace](https://www.evertrace.io/) | Europe-focused | ## Later Stage Signal Providers Typically, these providers will be better for Series A+ investors tracking growth signals and company momentum. | Provider | Notes | | -------------------------------------- | ---------------------------------------------------------------- | | [Specter](https://www.tryspecter.com/) | Broad coverage. Easy to use data dumps from S3. | | [Harmonic](https://harmonic.ai/) | GraphQL API. Supportive of many events in the VC tech ecosystem. | Choosing between these is largely personal preference. Focus on how well you can filter to get signal from noise: that's what matters most with these tools. **Sit with your investment team to decide which one is best for you.** ## Considerations **Speed vs accuracy:** Signal providers prioritize speed over perfection. The data might have gaps or errors, but you're seeing it before anyone else. **Volume can be overwhelming:** These providers generate lots of signals. You need good filtering and scoring to avoid drowning in noise. **Complement, don't replace:** Signal data works best alongside traditional company and people data, not instead of it. See [Considerations](/guide/part-3-technical-foundations/data-providers/considerations) for cost and vendor relationship guidance. # Data Starter Kits Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-providers/starter-kits Recommended data stacks for different fund types, stages, and focus areas. ## Overview What data you need depends on your fund: stage focus, sector specialization, team size, and budget. A pre-seed fund sourcing emerging founders needs different data than a growth fund doing due diligence on Series B companies. This page outlines starter kits for different fund profiles. These are starting points, not prescriptions. Your specific needs will vary. ## Pre-Seed / Seed Focus You're looking for companies before anyone else knows about them. Signal data based on government registries matters more than comprehensive funding history, though even if you have a comprehensive funding database, you may not find all the companies you're interested in. **What you need:** * Data to support your macro and market trend research * Early-stage signal data (who's starting companies, what's trending) * Founder and team data (background, previous experience) * Basic company data (to track what you find) **What you probably don't need yet:** * Comprehensive funding databases (most of your targets won't be in them) * Detailed financial data (too early for meaningful financials) **Typical stack:** | Category | Recommendation | | ----------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | Early Signal data | [Gravity](https://www.gravity.inc/) (US) or [Evertrace](https://www.evertrace.io/) (Europe) | | People and Company data | Choose either [People Data Labs](https://www.peopledatalabs.com/) or [Coresignal](https://coresignal.com/) for coverage of people and companies | | Research tools | [Perplexity API](https://docs.perplexity.ai) for quick market research | At this stage, signal and people data matter more than comprehensive funding databases. Focus your budget there. *** ## Series A / Series B Focus You're evaluating companies with some traction. Need a balance of signal data and comprehensive coverage. **What you need:** * Funding history and investor data * Growth signals (hiring, web traffic, product launches) * Team composition and changes * Competitive landscape data **What you probably don't need:** * Deep public market data * Heavy patent/research databases (unless sector-specific) **Typical stack:** | Category | Recommendation | | ------------ | --------------------------------------------------------------------------------------------------------------------------------------------------- | | Company data | [Crunchbase](https://www.crunchbase.com/) or [Dealroom](https://dealroom.co/), if you can splurge: [PitchBook](https://www.pitchbook.com/) | | Signal data | [Specter](https://www.tryspecter.com/) or [Harmonic](https://harmonic.ai/) for growth signals | | People data | [People Data Labs](https://www.peopledatalabs.com/), [Coresignal](https://coresignal.com/), or [MixRank](https://mixrank.com/) for team composition | | Web traffic | [SimilarWeb](https://www.similarweb.com/) if evaluating consumer companies | | Research | [Perplexity API](https://docs.perplexity.ai) or [Exa](https://exa.ai/) for competitive research | This is the "balanced" tier. You need both signal data (to find companies with momentum) and comprehensive company data (for due diligence). This is where you might start to experiment with "flat files" instead of APIs (see [Accessing Data](/guide/part-3-technical-foundations/data-providers/accessing-data)) *** ## Growth / Late Stage Focus You're doing deeper due diligence on established companies. Comprehensive data and financial metrics matter most. **What you need:** * Comprehensive funding databases * Financial and operational metrics * Market and competitive analysis * Public company comparables **What you probably don't need:** * Early-stage signal data (your targets are already known) **Typical stack:** | Category | Recommendation | | -------------- | ------------------------------------------------------------------------------------------------------------------- | | Company data | [PitchBook](https://pitchbook.com/) for comprehensive financials, valuations, cap tables | | Financial data | [S\&P Capital IQ](https://www.spglobal.com/marketintelligence/en/solutions/sp-capital-iq-platform) for public comps | | People data | [People Data Labs](https://www.peopledatalabs.com/) for team composition and hiring trends | At this stage, the coverage and quality of premium data becomes worth the investment. You need detailed financials, valuation history, and deal terms that lighter providers don't offer. You probably also need full data dumps, rather than just API access. *** ## Deep Tech / Bio Focus You're evaluating technical founders and novel technology. Research and patent data become critical. **What you need:** * Academic publication databases * Patent and IP data * Technical founder backgrounds * Research institution connections **Additional considerations:** * Many deep tech companies won't appear in standard funding databases until later * Founder evaluation requires different signals (publications, citations, lab affiliations) **Typical stack:** | Category | Recommendation | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Academic | [arXiv](https://arxiv.org/) (AI/ML, physics), [PubMed](https://pubmed.ncbi.nlm.nih.gov/) (bio/healthcare) | | Research tools | [Semantic Scholar](https://www.semanticscholar.org/) for citation networks and research impact or [Lens](https://www.lens.org/) for linking patents to academic research | | People data | [People Data Labs](https://www.peopledatalabs.com/) for founder backgrounds | **For bio/healthcare specifically, add:** | Category | Recommendation | | --------------- | ------------------------------------------------------------------------------------------- | | Clinical trials | [ClinicalTrials.gov](https://clinicaltrials.gov/) | | Drug pipeline | [BioMedTracker](https://www.biomedtracker.com/) for pipeline intelligence | | FDA data | [FDA databases](https://www.fda.gov/drugs/drug-approvals-and-databases/drugsfda-data-files) | Research and patent data are mostly free. Your budget goes toward people data and specialized tools like BioMedTracker. *** ## Regional / Sector Specialist You focus on a specific geography or vertical. Niche data providers often have better coverage than generalists. **What you need:** * Regional/sector-specific data providers * Local market intelligence * Sector-specific signals and metrics **Key insight:** Generalist data providers often have weak coverage outside US tech. If you invest in Europe, Asia, or specific verticals, look for specialized providers who focus on your market. **By geography:** | Region | Recommendations | | ------ | ----------------------------------------------------------------------------------------------------------- | | US | [Gravity](https://www.gravity.inc/) for signals, [Crunchbase](https://www.crunchbase.com/) for company data | | Europe | [Evertrace](https://www.evertrace.io/) for signals, [Dealroom](https://dealroom.co/) for company data | **By sector:** | Sector | Recommendations | | ----------- | --------------------------------------------------------------------------------------------------------------------- | | Consumer | [SimilarWeb](https://www.similarweb.com/) for web traffic, [data.ai](https://www.data.ai/) for mobile apps | | E-commerce | [Jungle Scout](https://www.junglescout.com/) for Amazon, [SimilarWeb](https://www.similarweb.com/) for traffic | | Real estate | [CARTO](https://carto.com/) or [SafeGraph](https://safegraph.com/) for location intelligence | | Fintech | [SEC EDGAR](https://www.sec.gov/edgar) for filings, standard company providers for funding data | | Climate | [EIA](https://www.eia.gov/) for energy data, [EPA databases](https://www.epa.gov/enviro/data-downloads) for emissions | The key is finding providers with deep coverage in your specific market rather than relying on generalists. *** ## Budget Considerations Your data budget should scale with fund size and strategy: * **Small fund (under \$50M):** Focus on 1-2 core providers. Start with what you absolutely need. * **Mid-size fund (\$50-250M):** Can afford broader coverage. 3-5 providers typical. * **Large fund (over \$250M):** 3-5 providers (but typically more expensive), then additional budget reserved for project or deal specific data sources. # Data Quality and Validation Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-quality How to minimize data quality problems, surface uncertainty, and maintain trust with your investment team despite the inherent messiness of external data. ## Overview You will show bad data to a GP. It's inevitable. A company in your market map that shut down two years ago. Funding amounts that are wrong. Employee counts that are wildly inaccurate. A competitor analysis that misses obvious competitors or includes companies that aren't actually competitive. This happens to everyone building data infrastructure for VC. The question isn't whether you'll have data quality problems. The question is how you minimize them, how you surface them when they exist, and how you maintain trust with your investment team when bad data inevitably makes it through. Data quality in VC is particularly challenging because you're aggregating data from multiple external sources, none of which are perfect. Crunchbase has gaps. PitchBook has errors. LinkedIn data is stale. Your CRM has whatever your team bothered to enter. Company websites lie. Every data source has different coverage, accuracy, and freshness characteristics. Unlike operational systems where you control data entry and can enforce validation rules, VC data comes from the outside world. You don't control when companies update their information. You can't force data vendors to be more accurate. You're playing defense, trying to catch errors before they damage your credibility with the people who make investment decisions. This chapter covers how to think about data quality in VC, how to choose and validate data sources, how to handle conflicting information, and how to build trust with your investment team despite the inherent messiness of external data. ## Selecting and Evaluating Data Sources Not all data is equally trustworthy. Data quality starts with choosing the right sources and understanding what you can trust them for. **The trust hierarchy** **High trust:** Your firm's internal data (portfolio companies, investment amounts, ownership, LP data). When external sources conflict with your internal data, your data is almost always correct. Use it as ground truth. If it's incorrect, you have a workflow issue, not a technical issue. **Medium trust:** Established vendors for their core competency. PitchBook for funding data. CoreSignal for employee counts. Your CRM for companies you track. These are reliable enough for analysis, but understand the tradeoffs: PitchBook is accurate but slow. Signal providers like Specter or Harmonic are fast but incomplete. Pick based on your use case: accuracy for LP reporting, speed for deal flow alerts. **Low trust:** Scraped data, meeting transcriptions, AI summaries. Useful for context and exploration, not definitive analysis. Transcriptions mishear names. Scraped data is stale. AI hallucinates. **Validate with your portfolio** Take 10-20 portfolio companies and check vendor accuracy: funding data, employee counts, descriptions, update speed. This reveals which vendors are reliable for which data types. Use this analysis to pick your authoritative source for each data type. Don't guess. **Pick one source per use case** The best way to avoid data quality problems is to not merge sources unnecessarily. Analyzing total market funding? Pick PitchBook and use only that. Don't combine multiple sources, entity resolution creates new errors. Only merge when each source provides unique value and it's required for the type of analysis you're doing. ## The Validation Challenge You want validation rules that catch errors before bad data reaches your investment team. In theory, this makes sense. In practice, it's much harder than you'd expect. **Why automated validation is limited** In operational systems, you can validate that email addresses have @ symbols, that phone numbers are numeric, that dates are in the future for scheduled events. These rules work because the domain is constrained and you control data entry. VC data doesn't work like this. Should Seed Rounds be less than \$1B? Usually, but Thinking Machines raised \$2B right off the bat. Should valuations increase with each round? Usually, but down rounds happen. Should funding dates be chronological? Usually, but sometimes rounds are announced retroactively. Should employee counts increase over time? Usually, but companies do layoffs. Every validation rule you write will have exceptions. Real companies don't follow clean patterns. Markets are messy. You'll either have rules so strict they flag too many false positives (wasting time investigating legitimate data), or rules so loose they miss real errors. **Domain knowledge over automation** The most effective validation is humans with industry knowledge reviewing data. Someone who knows the fintech space will immediately spot that "Stripe raised a \$10M Series A in 2023" is wrong because Stripe raised their Series A over a decade ago. Someone who follows AI infrastructure companies will know which competitors belong in a market map. Automated systems won't catch these errors because they lack context. This is why **you should know the industries and markets your fund focuses on**. If you're building data infrastructure for a fintech fund, learn fintech. If you're supporting a deep tech fund, understand deep tech. Domain knowledge is your best validation tool. **The gut feel test** When you build an analysis, look at it yourself before showing it to anyone. Does it pass the gut feel test? If you're showing a market map of 200 AI companies and you've never heard of 150 of them, something is probably wrong with your sourcing criteria. If funding totals seem way too high or too low compared to what you know about the market, investigate. Your intuition catches problems that automated validation misses. This is another reason why domain knowledge matters. You need a gut feel for what's reasonable. ## Handling Data Staleness Data gets stale. Companies update their information irregularly. Data vendors scrape on their own schedules. Your CRM is only as fresh as your last conversation with a company. Staleness is inevitable, but you can manage it. **Surface freshness to users** Always show when data was last updated. If you're displaying employee count, show "500 employees (as of Dec 2024)." Don't present stale data as if it's current. This does two things. First, it sets appropriate expectations. Users know whether to trust the data for current decisions. Second, it shifts responsibility. If a GP uses 6-month-old employee counts and makes a decision based on that, they knew the data was old. You're not hiding staleness. **Push vendors to refresh** Some data vendors have refresh or rescrape endpoints. If you need current data on a specific company, you can request that they update it. This works for high-priority companies (companies you're actively diligencing, portfolio companies) but doesn't scale to refreshing your entire database daily. Use selective refresh strategically. When you're preparing a market map for Monday's partner meeting, refresh the key companies in that market over the weekend. Don't try to keep everything fresh all the time. **Know when staleness matters** For some use cases, stale data is fine. If you're analyzing funding trends over the past five years, using data from three months ago is perfectly adequate. If you're building a list of potential competitors for a portfolio company, employee counts from six months ago are probably close enough. Staleness matters most for real-time decision-making. If a GP is deciding whether to take a meeting with a company this week, you want current information. If you're doing background research on a market, last quarter's data is fine. Understand which of your use cases are time-sensitive and prioritize freshness there. ## Building Trust with GPs You will show bad data to your investment team. How you handle it determines whether they trust you going forward. **Demo to friendly associates first**. Before presenting to partners, show it to a junior associate who knows the market. They'll spot embarrassing mistakes before decision-makers see them. An associate telling you "Airbnb didn't raise a Series A in 2023" is much better than a GP saying it in front of the partnership. **Always cite sources**. "Funding data from PitchBook, as of January 2026." "Employee counts from LinkedIn." This sets expectations (people know PitchBook is reliable, LinkedIn is approximate) and lets you shift blame when data is wrong. "This is what PitchBook reported" is better than "this is what I analyzed." **Entity resolution makes blame-shifting harder**. When you merge multiple sources and the data is wrong, whose fault is it? The vendor, the matching algorithm, or your choice of which source to trust? This is a hidden cost of entity resolution. Single-source analyses let you point to the provider when things are wrong. **Learn from each mistake**. When bad data reaches GPs, figure out why and fix the root cause. Report errors to vendors. Improve entity resolution. Don't use low-trust data for high-stakes analysis. Each mistake is a chance to improve your infrastructure. **Accept imperfection**. Even with perfect processes, external data will have errors. Set this expectation with your team: "We use PitchBook, which is generally accurate but has occasional errors." Better than promising perfect data and failing to deliver. ## Prevention and Detection Good data quality comes from both preventing bad data from entering your systems and detecting it when it does. **Prevention: Good ingestion practices** When you ingest data from external sources: * Normalize data (consistent date formats, consistent naming, lowercase domains) * Store what you received so you can trace errors back to the source * Version your data so you can see what changed and when * Flag data from low-trust sources so it doesn't get used inappropriately These practices don't prevent all errors, but they prevent careless mistakes and make debugging easier when problems arise. **Detection: Monitoring for obvious errors** Set up automated tests that run frequently to catch data quality issues before they reach users. [dbt](https://www.getdbt.com/) tests are particularly good for this. They let you define tests in SQL that run against your data warehouse and alert you when things break. **Examples of useful dbt tests for VC data:** * Uniqueness: Each company should have only one canonical ID * Not null: Required fields like company name, founding date are populated * Relationships: Every investment references a valid company and investor Run these tests frequently (daily or after each data ingestion). When tests fail, you get alerts before bad data reaches dashboards or analyses. **Human review for important analyses** Anything that will be seen by GPs or partners should have human review. You look at it (gut feel test), a friendly associate looks at it (domain knowledge check), and only then does it go to decision-makers. This doesn't scale to every query, but for high-stakes deliverables, human review is essential. ## The Bottom Line Data quality in VC is hard. You're aggregating external data you don't control, from sources with different accuracy and freshness characteristics, for users who make high-stakes decisions based on that data. You will show bad data to GPs. It's inevitable. Your job is to minimize errors, surface uncertainty when it exists, and maintain trust despite the inherent messiness of external data. **Understand trust levels**: Internal data > established vendors > scraped data. Use this to guide what data you use for what purposes. **Use portfolio companies to validate**: Test vendor accuracy against companies you know well. Use this to pick authoritative sources for each data type. **Pick one source per problem**: Don't try to merge PitchBook and signal providers for the same use case. Pick the best source and stick with it. Entity resolution creates new opportunities for errors. **Domain knowledge over automation**: Automated validation rules have limited value. Humans with industry knowledge catch more errors than algorithms. **Surface freshness**: Always show when data was last updated. Don't present stale data as current. **Always cite sources**: Let users know where data came from so they can calibrate trust and so you can shift blame when vendors are wrong. **Demo to friendly associates first**: Catch embarrassing mistakes before they damage credibility with GPs. **Accept imperfection**: External data will have errors. Set expectations appropriately rather than promising perfect data. In the next chapter, we'll cover data warehousing and analytics: once you have data (imperfect as it is), how do you structure it, analyze it, and surface insights that GPs actually use? # Data Warehousing and Analytics Source: https://buildingfor.vc/guide/part-3-technical-foundations/data-warehousing When you need a data warehouse, which one to choose, how to structure your data using dbt, and how to use it for analytics GPs actually care about. ## Overview A data warehouse is a central repository where you store and analyze data from all your sources: CRM, data vendors, fund operations, portfolio companies, your research platform. Instead of querying different systems separately, you pull everything into one place where you can run analyses that span multiple data sources. For VC funds, a data warehouse serves multiple purposes. It's an audit log of what happened in your systems (who talked to which companies, when deals moved through your pipeline, how data changed over time). It's the foundation for analytics and dashboards. It's where you run ad-hoc queries when a GP asks "how many Series A companies in fintech raised in the last 6 months?" And it's where you prepare data for tools like LLMs or knowledge graphs. You don't need a data warehouse on day one. But it's a good second step after you have your CRM set up. Even if you're not immediately running sophisticated analyses, building the infrastructure early creates a foundation that becomes more valuable as your fund scales. This chapter covers when you actually need a data warehouse, which one to choose, how to structure your data (following dbt best practices), and how to use it for analytics that GPs actually care about. ## When You Actually Need a Data Warehouse Most funds don't start with a data warehouse. They start with a CRM, maybe some spreadsheets, and direct queries to data vendors. This works fine at first. But as you add more data sources, want to track changes over time, or need to run analyses that combine multiple sources, you hit the limits of this approach. **The right time to set up a data warehouse**: After your CRM is set up and being used consistently. The CRM is your source of truth for deal flow and relationships. A data warehouse extends that by: * Tracking historical changes (you can see how your pipeline evolved over time, not just current state) * Combining CRM data with external sources (merge your deal flow with funding data from PitchBook, etc.) * Running analyses that would be painful in your CRM's query interface * Serving as an audit log of activity (who talked to which companies, when deals moved stages, what data changed) Even if you're not immediately running complex analyses, having a data warehouse captures history. Six months from now, when a GP asks "how many companies did we talk to in Q3 and what happened to them?", you have the data. Without a warehouse, that history might be lost or scattered across systems. **Start simple**: Your first data warehouse might just be a daily snapshot of your CRM data. Load it into a database, maybe add some basic transformations, and you're done. As you add more data sources and more sophisticated analyses, the warehouse grows with you. But start with something simple that works rather than trying to build a comprehensive data platform from day one. ## Choosing Your Data Warehouse **For small funds: Start with Postgres**. If you're tracking tens of thousands of companies (not millions), Postgres works fine. It's simpler and cheaper than specialized warehouses, your team already knows it, and you can use the same database for applications and analytics. Before migrating to a specialized warehouse, consider Postgres extensions: **[pg\_mooncake](https://www.mooncake.dev/pgmooncake/)** (Iceberg/Delta Lake format on S3) or **[Hydra](https://github.com/hydradatabase/columnar)** (columnar storage, 200x+ speedups for analytics). These let you keep Postgres while getting analytics database benefits. **For larger funds: Snowflake, BigQuery, or Redshift**. When Postgres performance degrades (millions of companies, complex aggregations), migrate to a specialized warehouse. [Snowflake](https://www.snowflake.com/en/) (industry standard, cloud-agnostic, expensive), [BigQuery](https://cloud.google.com/bigquery) (fast, Google ecosystem, query-based pricing), or [Redshift](https://aws.amazon.com/redshift/) (AWS integration, cheapest if committed to AWS). Choose based on your cloud ecosystem. **Migration path**: Start with Postgres. Migrate when query performance becomes painful. The migration is straightforward - you're already structuring data in a warehouse pattern, just moving to a more powerful database. ## The ELT Approach with dbt The modern approach is ELT (Extract, Load, Transform) using [dbt](https://www.getdbt.com/). Traditional ETL transforms data before loading it. ELT loads raw data first, then transforms it using SQL inside the warehouse. This is simpler because all transformation logic is in SQL, version-controlled, and tested like application code. **How it works**: Extract data from sources (CRM, PitchBook, LinkedIn) and load it raw into your warehouse using tools like [Fivetran](https://www.fivetran.com/) or [Airbyte](https://airbyte.com/). Write dbt models (SELECT statements) to transform raw data: clean and normalize, combine sources, calculate metrics, build final tables (marts) for analysis. dbt handles orchestration, testing, and documentation automatically. Analysts query the transformed tables, not raw data. ```mermaid theme={null} flowchart LR subgraph Sources direction TB CRM[CRM] ~~~ CO[Company Data] ~~~ PPL[People Data] ~~~ SIG[Signal Data] ~~~ MTG[Meeting Notes] ~~~ FO[Fund Ops] end EL[Extract & Load] subgraph Warehouse RAW[(Raw)] STG[Staging] INT[Intermediate] MARTS[Marts] end subgraph Destination direction TB DASH[Dashboards] ~~~ SQL[Ad-hoc SQL] ~~~ MCP[MCP Servers] ~~~ APP[Apps] end Sources --> EL --> RAW RAW --> STG --> INT --> MARTS MARTS --> Destination ``` ## Organizing Your Data: dbt Best Practices The right way to structure data in your warehouse follows [dbt best practices](https://docs.getdbt.com/best-practices). Rather than reinvent this, follow their guide. Here are the key principles: **Layered structure**: Staging → Intermediate → Marts * **Staging**: Raw data from sources, lightly cleaned (data types fixed, column names standardized, but otherwise unchanged). One staging model per source table. These models just make raw data easier to work with downstream. * **Intermediate**: Business logic and transformations. Join data from multiple sources, calculate derived fields, handle entity resolution (linking records across sources). These are building blocks that other models reference. * **Marts**: Final tables optimized for specific use cases or business domains. These are what dashboards and analyses query. Examples: `companies_enriched` (companies with all their data from multiple sources), `portfolio_performance` (portfolio company metrics over time), `deal_flow_metrics` (pipeline analytics). **Why this structure works**: Staging isolates you from source changes (if PitchBook changes their schema, you fix the staging model, not every downstream model). Intermediate models are reusable (if multiple analyses need company funding data enriched from PitchBook, build it once in intermediate). Marts are optimized for consumption (denormalized, pre-aggregated, easy to query). **Incremental models**: For large tables with historical data (every deal flow snapshot, every change in company data), use incremental models that only process new or changed data rather than rebuilding the entire table. This keeps transformations fast as data grows. **Tests**: Define tests in dbt for data quality (uniqueness, not null, relationships between tables, custom validations). These run automatically and alert you when data breaks, similar to what we covered in [Data Quality](/guide/part-3-technical-foundations/data-quality) on data quality. **Documentation**: Document your models, columns, and business logic in dbt. It generates a website showing your entire data model, lineage graphs (which models depend on which), and descriptions of what everything means. This is essential as your warehouse grows and new people need to understand the data. **Follow the guide**: The [dbt best practices](https://docs.getdbt.com/best-practices) guide goes into much more depth on how to structure projects, name things, write modular SQL, and optimize performance. Follow it. dbt has become the standard for a reason, and following their patterns makes your code easier for others (or future you) to understand. ## Analytics and Dashboards The value of a data warehouse isn't just storing data. It's enabling analysis that helps your fund make better decisions. **What GPs actually use is very fund-specific** Different funds care about different metrics. A pre-seed fund might focus on pipeline metrics (how many companies talked to, conversion rates by source, time from first contact to investment). A growth equity fund might focus on portfolio performance (company growth rates, follow-on opportunities, markups/markdowns). A thesis-driven fund might focus on market coverage (how many companies tracked in each thesis area, market maps by sector). There's no universal dashboard that every fund needs. What matters is having data structured so you can quickly answer questions that come up. **The value is ad-hoc analysis** The biggest benefit of having a data warehouse isn't pre-built dashboards. It's being able to answer questions quickly when GPs ask them. "How many companies did we pass on in 2023 that later raised Series B from our target co-investors?" "Which of our portfolio companies are hiring in engineering?" "What's our average time from first meeting to investment decision?" With data in a warehouse, these queries are straightforward SQL. Without it, they require manually pulling data from multiple systems, matching companies, and piecing together answers. The warehouse makes ad-hoc analysis feasible. **Tools like Hex for quick dashboards** When you do need dashboards, tools like [Hex](https://hex.tech/) or [Mode](https://mode.com/) let you write SQL queries against your warehouse and turn them into interactive visualizations. These are much faster to build than custom dashboards, and they're easy to iterate on and share with the team. At Inflection, we used Hex. When a GP wanted to see something, we could spin up a dashboard in an hour or two: write SQL queries against the warehouse, add visualizations, share a link. This responsiveness is valuable. You're not waiting weeks for data engineering to build custom dashboards. You're answering questions as they come up. **Examples of common analyses**: * Pipeline metrics (deals in each stage, conversion rates, time in pipeline, source of deals) * Market maps (companies in specific sectors, with funding, stage, growth data) * Portfolio tracking (company metrics over time, which companies need attention) * Investor analysis (which co-investors appear frequently, warm intro paths to founders) These aren't universal. What your fund analyzes depends on your strategy, stage, and what questions your investment team cares about. The warehouse makes it possible to answer whatever questions come up. ## Common Mistakes Most mistakes in data warehousing come from either over-complicating things early or not following established patterns. **Over-complicating early**: Don't build sophisticated incremental models, complex transformations, and elaborate marts on day one. Start with basic staging models that load raw data, maybe a few simple transformations, and iterate from there. Add complexity only when you need it. **Not following dbt patterns**: If you're using dbt, follow their best practices. Don't invent your own project structure, naming conventions, or transformation patterns. The value of dbt is that it's a standard. When someone new joins or you need help, they can understand your code because it follows known patterns. **Choosing the wrong warehouse for your scale**: Snowflake for a 3-person fund is probably overkill. Postgres for a large fund with billions of rows is probably inadequate. Match your tool to your scale. Start simple (Postgres), graduate to specialized warehouses (Snowflake/BigQuery) when you need them. **Not capturing history**: If you're only storing current state (latest CRM snapshot), you lose the ability to analyze how things changed over time. Capture history from the start. This doesn't need to be sophisticated - even daily snapshots of your CRM let you reconstruct what your pipeline looked like at any point. **Building dashboards nobody uses**: Don't spend weeks building elaborate dashboards until you validate that people will use them. Start with ad-hoc SQL queries. When you find yourself running the same query repeatedly, that's when to build a dashboard. Let demand drive what you build. ## The Bottom Line A data warehouse is good infrastructure to set up after your CRM is working. Even if you're not immediately running sophisticated analyses, it captures history and creates a foundation for analytics as your fund grows. Use dbt for transformations. Load raw data into your warehouse, transform it with dbt models following their best practices. This is simpler and more maintainable than custom ETL pipelines. Follow the [dbt best practices guide](https://docs.getdbt.com/best-practices) for organizing your data. The value of a warehouse is enabling ad-hoc analysis. You can quickly answer questions when GPs ask them, spin up dashboards with tools like Hex, and run analyses that combine multiple data sources. What specific analyses matter is fund-specific. What matters is having the infrastructure to answer questions quickly. Don't over-complicate. Start with basic models, add complexity only when needed, and follow established patterns. The warehouse should make analysis easier, not create a massive engineering project. In the next chapter, we'll cover knowledge graphs: when relationship-heavy queries matter for market mapping, competitor analysis, and network effects, and how to approach implementation. # Emerging Trends Source: https://buildingfor.vc/guide/part-3-technical-foundations/emerging-trends LLMs, MCP for connecting AI to internal data, agent orchestration, file-native agents, and what hype to ignore. ## Overview Large language models are the most significant technology shift for VC infrastructure in the past decade. They've changed how you build tools, how you extract information from documents, how you analyze companies, and how fast you can ship features as a solo developer. If you're building VC technology in without using LLMs, you're working at a significant disadvantage. But LLMs are becoming baseline, not cutting-edge. Every fund is using ChatGPT. Every engineer is using Claude Code or Cursor. The competitive advantage isn't that you're using LLMs, it's how you're using them and what you're building on top of them. This chapter covers emerging trends that matter for VC technology: connecting AI to your internal data through MCP, the evolution from single coding sessions to orchestrated agent workflows, why file-native agents are replacing RAG systems, and what hype to ignore. ## MCP: Connecting AI to Your Internal Data (or Just Use CLI Tools) Model Context Protocol (MCP) is Anthropic's standard for connecting AI assistants to data sources. Instead of copy-pasting data into Claude or building custom integrations for every tool, you build MCP servers that expose your data in a standardized way. Claude Code (and other MCP-compatible tools) can then query your internal systems directly. **That's the official story. Here's the alternative take:** Some developers argue MCPs are unnecessary abstraction and you should just write CLI tools instead. Claude Code can already execute bash commands and call CLI tools. A simple CLI script that queries your database or CRM is more universal than an MCP server (works with any tool that can run bash), simpler to build (no SDK required), and easier to maintain. The debate is ongoing, and it's not clear which approach will win. For now, here's practical guidance: if you're building something simple (query your database, fetch data from your CRM), start with a CLI tool. If you need features MCP provides (structured resources, interactive prompts, complex tooling), build an MCP server. Both approaches work for connecting AI to your internal data. **Why this matters for VC funds** You have valuable data scattered across systems: companies in your CRM, research in your data warehouse, portfolio metrics in various dashboards, memos in Google Docs or Notion. When you're building features or analyzing data, you currently need to manually pull information from each system, paste it into prompts, and context-switch constantly. MCP servers let AI tools access this data directly. You can ask Claude Code "show me all Series A companies in fintech we've talked to in the last 6 months" and it queries your CRM through an MCP server. You can ask "what are the latest metrics for our portfolio companies?" and it pulls from your data warehouse. The AI has the same access to data that you do, without you needing to be the intermediary. **The broader trend: Connecting AI to internal data** Whether through MCP servers, CLI tools, or other approaches, the key insight is that AI tools become dramatically more useful when they can access your internal data directly. Instead of being a general-purpose assistant, they become specialized tools that understand your fund's portfolio, pipeline, and research. This could mean: * Querying your CRM to find companies or check deal status * Running SQL against your data warehouse to analyze portfolio performance * Searching through investment memos and research to find relevant context * Pulling metrics from portfolio company dashboards The mechanism (MCP vs. CLI vs. something else) matters less than the outcome: AI tools that can answer questions about your fund's specific data without you manually feeding them information. **LLM providers moving up the stack** This trend represents LLM providers (Anthropic, OpenAI, etc.) moving beyond just providing model APIs to building full development environments with data access. Claude Code isn't just a better coding assistant - it's becoming a platform for building and using internal tools. The new [Claude Cowork](https://support.claude.com/en/articles/13345190-getting-started-with-cowork) and [Claude in Excel](https://support.claude.com/en/articles/12650343-claude-in-excel) are further examples of LLM providers moving into the application layer. The latest development is [MCP Apps](https://blog.modelcontextprotocol.io/posts/2026-01-26-mcp-apps/), which allow MCP servers to expose UI components directly to the LLM client. Instead of just providing tools that return text, an MCP server can now render interactive interfaces: forms, charts, tables, approval workflows. This blurs the line between "AI assistant" and "application platform" even further. Watch this space. The tooling will evolve, standards may change, but the direction is clear: AI tools will increasingly integrate with your internal systems rather than operating in isolation. ## AI-Assisted Development: From Sessions to Orchestration [Claude Code](https://code.claude.com/docs/en/overview) and [Cursor](https://cursor.com/) have already changed how you build software. You can implement features in hours that previously took days. You can build entire applications as a solo developer that previously required small teams. This is the current state, and it's already transformative. But we're at the beginning, not the end, of AI-assisted development. **Current state: Single session, single developer** Today, you start a Claude Code session, describe what you want to build, and Claude helps you write code, debug issues, and ship features. When the session ends (or you hit context limits), you start fresh. You're still fundamentally working alone, just with a very capable assistant. This is already powerful. As covered in [Choosing Your Stack](/guide/part-3-technical-foundations/choosing-your-stack), choosing popular technology stacks (Next.js, TypeScript) means AI coding tools work better and you ship faster. But there are limits: complex features that span multiple services, background work that takes hours to run, coordinating changes across many files. **Near future: Agent orchestration** The next evolution is multiple AI agents working in parallel on the same codebase through git worktrees. Instead of one Claude Code session, you might have ten agents simultaneously: * Agent 1 implements the frontend UI for a new feature * Agent 2 builds the backend API endpoints * Agent 3 writes tests for both * Agent 4 updates documentation * Agent 5 handles database migrations * Agents 6-10 work on related features or refactoring Each agent works in its own git worktree (a separate working directory pointing to a different branch). They can work independently without conflicts. When agents finish their work, they create PRs that you review and merge. The agents coordinate through the git repository: they see each other's changes, can pull updates, and understand the evolving codebase. This isn't science fiction. The building blocks exist: git worktrees are a standard git feature, Claude Code can already work with git, and orchestration systems are being built: * [Conductor](https://www.conductor.build/) * [Emdash](https://www.emdash.sh/) * [Superset](https://superset.sh/) **What this means for VC tech** Solo developers will be able to build and maintain even more ambitious systems. Maintaining multiple internal tools that currently requires your full attention becomes more manageable when agents handle routine updates and testing. You don't need to do anything now except be aware this is coming. When agent orchestration tools mature, the same principles from [Choosing Your Stack](/guide/part-3-technical-foundations/choosing-your-stack) apply: use boring, proven technology that AI tools understand. Use TypeScript and Next.js. Structure your code clearly. Write good documentation. These practices make both current AI tools and future agent orchestration more effective. **Personal AI agents: Beyond development** The same pattern is emerging for personal productivity. Tools like [Moltbot](https://molt.bot/) ([formerly Clawdbot](https://x.com/moltbot/status/2016058924403753024), open source, runs locally) let you interact with an AI assistant through WhatsApp, Telegram, Slack, or iMessage. The agent has full system access: it can read files, execute commands, manage your calendar, send emails, and automate workflows across 50+ integrations. This is worth watching as inspiration for what autonomous agents can do. But approach with caution: giving an agent full system access has real security implications. Understand what you're installing before running it on a machine with access to fund data. **Sandboxed environments for agent execution** As agents gain more autonomy, sandboxing becomes critical. Running AI-generated code or giving agents system access on your local machine is risky. A new category of infrastructure is emerging to address this: isolated execution environments purpose-built for agents. * [Modal Sandboxes](https://modal.com/products/sandboxes): Container-based execution with sub-second startup, network controls for restricting outbound access, and scaling to 50,000+ concurrent sandboxes. * [Vercel Sandbox](https://vercel.com/docs/vercel-sandbox): Ephemeral Linux VMs designed for AI agents and code generation, with SDK-first integration. * [Sprites](https://sprites.dev/): Hardware-isolated Firecracker VMs with persistent state, checkpointing, and layer 3 network policies. These platforms let you run untrusted code without exposing your production systems or local machine. The key features for security are network isolation (preventing data exfiltration), ephemeral environments (no persistent access), and resource limits (preventing runaway processes). If you're building agent-powered features that execute code or access external systems, consider running them in sandboxed environments rather than on your local machine or production infrastructure. ## File-Native Agents: Beyond RAG and Knowledge Graphs For the past two years, the standard approach to helping AI systems work with large document collections has been RAG (Retrieval Augmented Generation): chunk documents into pieces, embed them as vectors, store in a vector database, retrieve relevant chunks based on query similarity, stuff them into context. This approach is becoming obsolete. **Why RAG was necessary** RAG existed as a workaround for limited context windows. If you could only fit 8K or 32K tokens into context, you couldn't give an AI access to hundreds of documents. So you chunked documents, embedded them, and retrieved only the most relevant pieces for each query. Knowledge graphs were a similar workaround: extract entities and relationships, build a structured graph, query it to find relevant information. Both approaches created intermediate representations (vectors, graphs) because we couldn't work with documents directly. **What changed** Context windows are now large enough and getting larger. AI agents have file system access and can use tools (grep, specialized readers). Context compaction techniques let agents maintain understanding across indefinitely long sessions. This means agents can work with files directly, like humans do. No chunking, no embedding, no graph extraction. Just "here's a folder of investment memos, analyze the fintech companies we've evaluated." **File-native agents: Documents stay as documents** Instead of transforming documents into vectors or graphs, file-native agents: 1. Have direct access to files in their native formats (PDFs, markdown, spreadsheets) 2. Use tools to search and analyze (grep for keywords, specialized readers for PDFs) 3. Maintain context through compaction (file system holds artifacts, compacted context holds insights) The key insight: files work because they're a shared abstraction. They're not optimal for agents or humans individually, but they're common ground both can navigate. This shared interface enables collaboration. If you create agent-only structures (vector databases, proprietary knowledge graphs), you break the collaborative aspect. **What this could mean for VC infrastructure** Building complex RAG systems for your internal documents could be a thing of the past: no need to extract entities from investment memos and build knowledge graphs. These are solutions to problems that no longer exist. Instead: * Store research as markdown files in git repositories * Give AI agents file system access to these repositories * Let agents use grep, read files, and search naturally * Focus on context compaction and session management, not retrieval algorithms If you're using a tool like Claude Code, it already works this way. It has file access, uses grep and other tools, and manages context effectively. You don't need to build additional infrastructure. **The one exception: External vendor data** For large external datasets (all companies from PitchBook, millions of records), traditional database queries are still appropriate. File-native agents are for documents and internal research, not for structured data at scale. Continue using your data warehouse and SQL for that use case (as covered in [Data Warehousing](/guide/part-3-technical-foundations/data-warehousing)). **What about existing RAG systems?** If you already built a RAG system for your investment memos or research, you don't need to immediately rip it out. But when you're building new features or reconsidering your architecture, consider file-native approaches. They're simpler, more maintainable, and work better with modern AI tools. ## What to Ignore **Don't fine-tune models**. Foundation models work fine for almost all VC use cases (extracting data, analyzing companies, answering questions). Fine-tuning requires training data, evaluation infrastructure, ongoing maintenance, and rarely produces better results for most VC tasks. **Don't add AI just to say you have AI**. Build features that solve actual problems. "AI-powered market maps" that are just LLM text aren't better than human-curated maps. Use AI where it creates real leverage: extracting information from documents, processing unstructured data, helping engineers build faster. Not for theater. **AGI timelines don't matter for your job**. Whether AGI arrives in 3 or 30 years doesn't change what you should build today. Build tools that work now and are maintainable by humans. **Model selection is simpler than you think**. We use Claude Sonnet for most tasks, Claude Opus for deep reasoning. Don't optimize for tiny cost differences. Pick an LLM provider, stick with it, and focus on building features. ## Staying Current Technology for VC infrastructure evolves quickly. What's cutting-edge today becomes baseline within months. Here's how to stay up to date without spending all your time chasing trends: **General tech communities** * **[X](https://x.com)**: Follow engineers building in the AI space, VC tech practitioners, and companies building tools for VCs. The signal-to-noise ratio is low, but you'll see emerging tools and approaches before they're widely adopted. * **[Hacker News](https://news.ycombinator.com)**: The Show HN section surfaces new tools and libraries. The comments often contain practical wisdom from people who've tried things in production. Good for understanding what's actually working versus what's just hype. **VC-specific resources** * **[Data Driven VC](https://datadrivenvc.io/)**: Community and resources specifically for people building data infrastructure at VC funds. Much better signal than general tech communities for VC-specific challenges. * **[Vestberry VC Day](https://vcday.vestberry.com/)**: Conferences focused on VC operations and technology. Good for understanding what funds of all sizes are building and what tools are emerging in the ecosystem. **The balance: Follow loosely, adopt carefully** Don't try to implement every new tool or technique you see. Most trends don't matter for your fund. Follow these resources to build context about what's possible and what direction the industry is moving, but only adopt new approaches when they solve actual problems you're experiencing. The goal isn't to use the latest technology. It's to build tools that help your fund invest better. Sometimes that means adopting new approaches early. More often it means sticking with proven technology and focusing on execution. ## The Bottom Line LLMs have fundamentally changed how you build VC technology, but they're becoming baseline rather than differentiating. The competitive advantage is how you integrate AI into your workflows, not that you're using AI. Connect AI tools to your internal data, whether through MCP servers, CLI tools, or other approaches. The mechanism matters less than the outcome: AI that can query your CRM, data warehouse, and research directly. Understand where AI-assisted development is heading: from single sessions to orchestrated agents working in parallel through git worktrees. Prepare by using technology stacks AI tools understand (TypeScript, Next.js, clear code structure) and writing good documentation. Skip building RAG systems for internal documents. File-native agents with large context windows and file system access work better. Store research as files, give agents file access, let them use tools naturally. Focus on context compaction, not retrieval algorithms. Ignore the hype: don't fine-tune models, don't add AI for theater, don't worry about AGI timelines, and don't overcomplicate model selection. Just use Claude Sonnet for most things and Opus when you need deeper reasoning. The next 6 months will bring better tooling for agent orchestration, more mature MCP ecosystems, and continued improvements in model capabilities. But the fundamentals won't change: use AI where it creates real leverage, integrate it into actual workflows, and focus on solving problems rather than using the latest technology for its own sake. # Entity Resolution Source: https://buildingfor.vc/guide/part-3-technical-foundations/entity-resolution Determining which records across different data sources refer to the same real-world entity - one of the hardest and most important data problems in VC. ## Overview When pulling company data from multiple sources (Crunchbase, PitchBook, LinkedIn, your CRM), the same company appears differently in each. "Acme Inc," "Acme Corporation," "Acme," "ACME, Inc.", all the same company, but your system doesn't know that. Entity resolution is determining which records refer to the same real-world entity. Without it, you can't aggregate data, build accurate analytics, or track companies properly. It's one of the hardest and most important data problems in VC. Getting it right was significant to EQT's Motherbrain. This chapter covers why it matters, how to approach it (start simple, build up only when necessary), and what it takes to build your own system. Spoiler: start with URL matching and avoid building a full system as long as possible. ## The Entity Resolution Problem **Company name variations**: "Stripe, Inc." vs "Stripe" vs "Stripe Payments" vs "Stripe Inc" (no period). These are all the same company, but string matching fails. Companies also rebrand ("Facebook" became "Meta"), get acquired (should "Instagram" be separate from Meta or merged?), and use different legal names in different jurisdictions. **URL variations**: "stripe.com" vs "[https://stripe.com](https://stripe.com)" vs "[www.stripe.com](http://www.stripe.com)" vs "stripe.com/". Same domain, different formats. Some sources include the protocol, others don't. Some include www, others don't. Trailing slashes are inconsistent. **Location variations**: "San Francisco, CA" vs "San Francisco, California" vs "SF" vs "San Francisco, CA, USA". Same location, different representations. Some sources use city and state, others add country, others use abbreviations. **People name variations**: "Alex Smith" vs "Alexander Smith" vs "A. Smith". Same person, different formats. People also change names (marriage, personal preference), use nicknames professionally, and have names that transliterate differently from other languages. **Data staleness**: A company moved offices, changed their URL, or was acquired. Some data sources have updated information, others don't. Which is correct? How do you handle temporal changes while maintaining entity consistency? **Missing data**: Not every source has every field. Crunchbase might have a LinkedIn URL, but PitchBook doesn't. One source has founding date, another doesn't. You need to match entities even with incomplete information. The fundamental challenge is that there's no universal, stable identifier for companies or people across all data sources. Each source assigns their own IDs. You need to create mappings between these IDs to know when records refer to the same entity. ## Why Entity Resolution Matters for VC Without entity resolution, your data infrastructure breaks in multiple ways. **Analytics are wrong**: If "Acme Inc" and "Acme Corporation" are both in your database as separate companies, your portfolio size is overstated. Your market maps show duplicate companies. Your metrics count the same funding round twice if it appears in both Crunchbase and PitchBook under slightly different names. **You can't aggregate data**: You want to combine Crunchbase's funding data with PitchBook's valuation data and LinkedIn's employee count for the same company. Without entity resolution, you can't merge these data sources. You're stuck with fragmented, incomplete information about each company. **Knowledge graphs break**: If the same company appears as multiple nodes in your graph, relationships are split incorrectly. Investors appear to have invested in fewer companies than they actually did. People appear to work at multiple companies simultaneously. Competitive relationships are missed. **Sourcing is inefficient**: Your sourcing tool identifies "Acme Inc" as a potential investment. But you already passed on this company three months ago when it appeared in your CRM as "Acme Corporation." Without entity resolution, you waste time re-evaluating the same companies. **Portfolio tracking fails**: You invested in a company, but your fund operations system, CRM, and data warehouse all have different records for this company with no linkage. When the company raises a Series B (triggering a markup), you have to manually update multiple systems because they can't automatically recognize it's the same company. This isn't a nice-to-have. Entity resolution is foundational to making any cross-source data system work correctly. ## Start Simple: See How Far You Can Get Before building anything sophisticated, see how far simple matching rules take you. **Match on URLs**: Company websites are relatively unique and stable. If two records have the same domain (stripe.com), they're almost certainly the same company. Normalize the URLs first (remove protocol, remove www, remove trailing slashes, lowercase everything), then compare. This catches most matches. **Match on LinkedIn URLs**: LinkedIn company pages have stable identifiers in their URLs (linkedin.com/company/stripe). If two records share the same LinkedIn URL, they're the same company. Same approach: normalize then compare. This also works well for people data. **Match on email addressses**: If you don't have reliable LinkedIn URLs for people, email addresses are usually the next best thing. **Exact string matches**: After normalizing names (lowercase, remove punctuation, standardize "Inc" vs "Incorporated"), exact string matches catch many cases. "Stripe Inc" and "Stripe, Inc." become "stripe inc" and match. These simple rules will catch 70-80% of matches without any sophisticated infrastructure. Implement these first. Only move to more complex solutions when simple matching stops working. **Why this matters**: Building an entity resolution system is expensive in engineering time and ongoing maintenance. Every hour spent on entity resolution is an hour not spent on research platforms, deal flow tools, or portfolio support. If simple rules handle most cases, that might be enough for your fund's scale. ## Sometimes You Can Avoid Entity Resolution Entirely Before investing in sophisticated entity resolution, consider whether you actually need to combine data sources for your use case. **Use single authoritative sources**: Instead of stitching together funding data from Crunchbase, PitchBook, and your CRM, pick the best source for funding data (probably PitchBook) and use only that for market analysis. Don't try to combine them. This avoids the entity resolution problem entirely for that use case. **Examples where single-source works**: * Analyzing total funding in a market? Use PitchBook exclusively. * Tracking employee counts? Use LinkedIn data from one of the [People Data Providers](/guide/part-3-technical-foundations/data-providers/people-data) only. * Portfolio valuations? Use your fund operations system as the single source of truth. **When you actually need entity resolution**: Cross-source analysis where each source provides different valuable data. For example, Crunchbase has good early-stage company coverage, PitchBook has better late-stage data, and People Data Labs has employee information. If you need all three types of data for the same company, you need entity resolution to link them. But if you can answer your question with just one source, don't create unnecessary complexity. **The pragmatic approach**: Start by identifying which questions you're trying to answer. For each question, can you answer it with a single data source? If yes, use that source exclusively for that use case. Only invest in entity resolution when you have validated use cases that genuinely require combining multiple sources. This saves enormous engineering effort. **Entity resolution is hard.** Avoiding it when possible is smart, not lazy. ## Entity Resolution as a Service Before building your own system, consider whether services exist that solve this problem for you. **Services available**: Several vendors offer entity resolution as a service: * **[Senzing](https://senzing.com/)**: Real-time entity resolution focused on messy people and organization data. Available on AWS Marketplace, deploys in your infrastructure for data privacy and compliance. Good for company matching with inconsistent names and addresses. * **[Tamr](https://www.tamr.com/entity-resolution)**: AI-native entity resolution with specific B2B capabilities. Offers automated matching and supports Dun & Bradstreet enrichment for company data. Focuses on master data management and customer 360 use cases. * **[AWS Entity Resolution](https://aws.amazon.com/entity-resolution/)**: Amazon's native service with flexible, configurable workflows. Integrates directly with other AWS services if you're already in that ecosystem. These services maintain their own canonical entity IDs and provide mappings from various data sources to their IDs. **The cost problem**: These services are tremendously expensive. Enterprise pricing often starts at tens of thousands per year and scales with usage. For large funds managing extensive data infrastructure, this might be justified. For smaller funds or those just starting to build data capabilities, the cost is prohibitive. **When it makes sense**: If you're at significant scale (tracking hundreds of thousands of companies, integrating 5+ data sources, have engineering resources but want to avoid building matching infrastructure), a service might make sense. If you're smaller or earlier, the cost probably doesn't justify the benefit compared to simpler approaches. **Vendor lock-in considerations**: Once you adopt a service's entity IDs, migrating to a different system is painful. All your internal systems reference these IDs. Switching providers means remapping everything. Make sure the vendor is reliable and will be around long-term. Consider services, understand their costs, but don't assume they're the only solution. Many successful funds built their own entity resolution because services didn't exist or were too expensive. ## Building Your Own Entity Resolution Service When simple matching stops working and services are too expensive, you need to build your own system. This is where it gets hard. **The canonical ID approach**: Assign one internal ID per entity. Every company or person gets exactly one canonical ID in your system. All records from external sources (Crunchbase, PitchBook, LinkedIn, your CRM) map to these canonical IDs. When you query for a company, you query by canonical ID, which pulls data from all sources that reference that entity. This canonical ID becomes your "golden record" or "single source of truth" for each entity. External IDs come and go, but your canonical ID is stable. **Matching strategies for companies**: When comparing two company records to decide if they're the same entity, match on multiple fields: * **Name**: Fuzzy string matching with normalization. Convert to lowercase, remove punctuation, standardize abbreviations ("Inc", "Corp", "Ltd"). Use string similarity algorithms (Levenshtein distance, Jaro-Winkler) to catch typos and minor variations. * **URL/Domain**: As mentioned earlier, this is your strongest signal. If domains match, very high confidence it's the same company. * **LinkedIn URL**: Similarly strong signal if available. * **Location**: City and country provide supporting evidence. "San Francisco" and "SF" should normalize to the same value. Location alone isn't sufficient (multiple companies in the same city), but combined with name similarity it strengthens confidence. * **Founding date**: If both records have founding dates and they're close (within a year), that supports a match. But founding dates are often missing or wrong in data sources, so don't rely on this. * **Description similarity**: Company descriptions can be compared using embeddings or keyword analysis. If two companies with similar names also have similar descriptions, higher confidence they're the same entity. **Probabilistic matching**: Instead of requiring exact matches on all fields, use a scoring system. If name is 90% similar, domain matches, and location matches, that's high confidence (>95%) they're the same entity. If only name is 70% similar and location matches, that's medium confidence (60-80%), might need human review. Define thresholds for automatic matching, human review, and rejection. **Matching strategies for people**: * **LinkedIn URL**: Strongest signal. LinkedIn profiles have stable IDs. * **Email**: If you have email addresses from multiple sources, exact match is strong signal. * **Name + Company**: If first name, last name, and current company all match, very likely the same person. **The incremental matching workflow**: When a new company record arrives from a data source: 1. Check if it has a URL or LinkedIn URL that matches an existing canonical entity → if yes, link it 2. If no URL match, check for high-confidence name + location match → if yes, link it 3. If medium confidence match, flag for human review 4. If no match found, create a new canonical entity This workflow runs every time you ingest data from external sources. ## Human-in-the-Loop: You'll Need It Automated matching will never be 100% accurate. You will have false positives (incorrectly merged entities) and false negatives (incorrectly split entities). You need tooling for humans to review and correct mistakes. **Review interface**: Show suggested matches with confidence scores. A human reviews the match (looking at names, URLs, descriptions, and any other data) and confirms or rejects it. This is especially important for medium-confidence matches where the algorithm isn't sure. **Merge interface**: When you discover two canonical entities that should be one (the algorithm missed a match), you need to merge them. This means: * Choosing which canonical ID to keep * Mapping all external references from the old ID to the new ID * Combining data from both entities into a single record * Preserving history (you want to know these were merged in case you need to undo it) **Split interface**: When you discover one canonical entity that should be two (the algorithm incorrectly merged different companies), you need to split them: * Create a new canonical ID for the second entity * Decide which external references belong to which entity (this is hard - you're untangling merged data) * Recalculate any analytics or relationships that were affected by the incorrect merge **Why you need these interfaces**: You will make mistakes. Automated matching will incorrectly merge "Acme Inc" (a fintech company in SF) with "Acme Corp" (a logistics company in NYC) because the names are similar and the algorithm wasn't confident enough to reject. You'll need to split them. Or the algorithm will miss that "Facebook" and "Meta" are the same company because they have different names and URLs. You'll need to merge them. These aren't edge cases. At scale, with hundreds of thousands of entities, you'll do splits and merges regularly. Build the tooling to make it easy, ideally for non-technical team members who understand the domain. ## Common Pitfalls **Matching too aggressively**: False positives are painful. If you incorrectly merge two different companies, all their data gets mixed together. Investors, employees, funding rounds, relationships: all incorrectly attributed to a single entity. This corrupts your analytics and is hard to untangle. Be conservative with automatic matching. When in doubt, flag for human review. **Matching too conservatively**: False negatives are also painful. If you don't merge the same company across sources, you have duplicate entities, fragmented data, and wrong analytics. Finding the right balance between false positives and false negatives is the art of entity resolution. **Not handling name changes over time**: Companies rebrand. Facebook became Meta. Google became Alphabet (with Google as a subsidiary). Your entity resolution needs to handle temporal changes. The same canonical entity might have multiple names over its history. Store name changes with effective dates so you know what to call the company at different points in time. **Not handling acquisitions properly**: Is Instagram a separate entity from Meta? For some purposes yes (you want to track Instagram as a product), for others no (financially it's part of Meta). You might need different entity models for different use cases: legal entities (separate companies) vs. operational entities (products within a parent company). **No interface for fixing mistakes**: If the only way to merge or split entities is by manually editing database records, your non-technical team can't help. Mistakes will accumulate because fixing them is too hard. Build interfaces that make corrections easy. **Assuming automated matching is enough**: Even with sophisticated algorithms, human review is necessary. Budget time for reviewing medium-confidence matches and investigating reported issues. This is ongoing operational work, not one-time setup. **Not preserving history**: When you merge or split entities, log what happened and when. You might need to undo a merge. You might need to understand why analytics changed after an entity resolution update. Audit trails matter. ## The Bottom Line Entity resolution is incredibly difficult. It's also fundamental to making multi-source data systems work correctly. Without it, you have duplicates, fragmented data, and wrong analytics. Start simple. Match on URLs and LinkedIn URLs. Normalize strings and do exact matching. See how far these basic rules take you. For many funds, especially smaller ones, this might be sufficient. Entity resolution services exist, but they're tremendously expensive. Consider them if you're at significant scale, but understand the costs and vendor lock-in implications. When you need to build your own system, use the canonical ID approach: one ID per entity, all external data maps to it. Match on multiple fields (name, URL, location, LinkedIn). Use probabilistic scoring to determine confidence. Automatically merge high-confidence matches, flag medium-confidence for human review. Build interfaces for splitting and merging entities. You will need them. Mistakes happen at scale, and you need tooling to fix them easily. Entity resolution was part of the secret sauce that made EQT's Motherbrain powerful. It's not a solved problem. It requires continuous investment and refinement. Don't underestimate how hard this is, but also don't over-engineer it before you need sophisticated matching. In the next chapter, we'll cover data quality and validation, which builds on entity resolution. Once you have canonical entities, you need to ensure the data about those entities is accurate, complete, and trustworthy. # Integrations Source: https://buildingfor.vc/guide/part-3-technical-foundations/integrations-and-apis Common integration patterns, validating API responses, handling webhooks and rate limits, and practical approaches to connecting VC tools. ## Overview Much of the technical work at a VC fund is building glue between tools. Connecting meeting notes to your CRM, data providers to your warehouse, portfolio data to dashboards. Some integrations exist out of the box, but many require custom code. This chapter covers common integration patterns, validation strategies, webhook handling, and rate limits. ## Common Integration Patterns The integrations you build fall into a few common patterns. **Tool-to-tool glue**: Connecting SaaS tools your team uses. Examples: * Meeting transcription tool (Granola, Otter) → CRM (Attio) to automatically log conversations * CRM → data warehouse for analysis * Email → CRM to track outreach * Calendar → CRM to log meetings Some of these have out-of-box integrations. Many require custom code to map fields, handle authentication, and deal with differences in data models. **Data vendor → your systems**: Pulling data from external providers and loading it into your infrastructure. These integrations usually run on schedules (nightly data syncs) or in response to specific events (when you add a company to your CRM, enrich it with PitchBook data). See [Accessing Data](/guide/part-3-technical-foundations/data-providers/accessing-data) for how vendors deliver data. **Internal service integrations**: If you're building multiple services (research platform, sourcing tool, internal APIs), they need to talk to each other. Your research platform might need portfolio company data from your data warehouse. Your sourcing tool might need to check your CRM to avoid suggesting companies you've already passed on. **LLM/AI integrations**: Many funds are building features that use LLMs. These require integrations with OpenAI, Anthropic, or other providers. Often combined with your own data (RAG systems pulling from your research or CRM). The common thread: you're moving data between systems, transforming it to fit different schemas, and handling failures when things break. ## Validating API Responses The biggest source of problems in integrations is trusting external APIs to return what you expect. API schemas change. Vendors return errors in unexpected formats. Required fields are sometimes null. Data types don't match documentation. **Never trust external APIs.** Validate everything. **Use validation libraries** For TypeScript: **[Zod](https://zod.dev/)**. For Python: **[Pydantic](https://docs.pydantic.dev/)**. These libraries let you define schemas for your data and automatically validate objects against them. ```typescript theme={null} // TypeScript with Zod import { z } from "zod" const CompanySchema = z.object({ id: z.string(), name: z.string(), founded_date: z.string().datetime().optional(), funding_total: z.number().positive().optional(), employee_count: z.number().int().positive().optional(), }) // When you get data from an API const response = await fetch("https://api.vendor.com/companies/123") const data = await response.json() // Validate it const company = CompanySchema.parse(data) // Throws if invalid // or const result = CompanySchema.safeParse(data) // Returns success/error if (!result.success) { console.error("Invalid data from API:", result.error) } ``` ```python theme={null} # Python with Pydantic from pydantic import BaseModel, field_validator from datetime import datetime from typing import Optional class Company(BaseModel): id: str name: str founded_date: Optional[datetime] = None funding_total: Optional[float] = None employee_count: Optional[int] = None @field_validator('funding_total') def funding_must_be_positive(cls, v): if v is not None and v < 0: raise ValueError('funding must be positive') return v # When you get data from an API response = requests.get('https://api.vendor.com/companies/123') data = response.json() # Validate it try: company = Company(**data) except ValidationError as e: print(f'Invalid data from API: {e}') ``` **Why this matters** Without validation, bad data silently flows into your system. A field that's supposed to be a number is suddenly a string. A required field is null. These errors cascade: your data warehouse has invalid data, your dashboards show wrong information, your analyses are incorrect. With validation, you catch errors at the boundary. When an API returns bad data, you know immediately. You can log the error, alert yourself, and handle it gracefully rather than letting corrupted data spread through your systems. **Validate both inbound and outbound data** When you're calling external APIs, validate the data you're sending. This catches mistakes in your code before they hit the vendor's API. When you're exposing APIs for internal use, validate inputs from callers. ## Webhook Handling Many vendors (especially CRM systems) provide webhooks: they call your HTTP endpoint when events happen (company updated, deal stage changed, meeting logged). This is more efficient than polling their API constantly. **Setting up webhooks** You need: * An HTTPS endpoint the vendor can reach (use [webhook.site](https://webhook.site/) for local development with dummy data) * To register your endpoint with the vendor (usually through their dashboard) * To handle webhook verification (vendors send a signature to prove the request came from them) **Key considerations** **Verify signatures**: Always verify that webhooks actually came from the vendor. They send a signature (usually HMAC of the body using a shared secret). Verify this before processing. **Handle idempotency**: Vendors may send the same webhook multiple times (network retries, their infrastructure issues). Make your webhook handler idempotent: processing the same event twice should be safe. Track event IDs you've seen and skip duplicates. **Return quickly**: Webhook handlers should return 200 OK within a few seconds. Don't do expensive processing in the handler itself. Accept the webhook, queue the work, and return success. Process asynchronously. ## Rate Limits and API Costs External APIs have rate limits. PitchBook might allow 100 requests per minute. People Data Labs might allow 500 requests per day. Exceed these and you get 429 errors or get blocked. **API costs** Beyond rate limits, many vendors charge per request or per entity returned. This changes how you think about API usage: * Only request data you actually need * Cache responses so you don't request the same data repeatedly * Batch requests when possible * Validate input before making API calls (don't waste money on requests that will fail) Monitor your API spending. Set up alerts when spending exceeds thresholds. **Always respect vendor limits** Don't try to work around rate limits by spinning up multiple API keys or using proxies. Vendors notice and will block you. **Exponential backoff** When you hit a rate limit or get a transient error, retry with exponential backoff: 1. First retry: wait 1 second 2. Second retry: wait 2 seconds 3. Third retry: wait 4 seconds 4. Fourth retry: wait 8 seconds 5. Give up after 5 attempts Add jitter (random variation) to prevent many clients from retrying simultaneously. ## Error Handling Not all errors should be retried. Some are permanent, others are transient. **Categorize errors** * **Retriable**: 429 (rate limit), 503 (service unavailable), 504 (timeout), network errors. Retry with backoff. * **Non-retriable**: 400 (bad request), 401 (unauthorized), 403 (forbidden), 404 (not found). Fix the problem or skip the request. * **Context-dependent**: 500 (server error) might be temporary or might indicate a vendor bug. Retry a few times, but not indefinitely. **Use transactions** When working with databases, wrap operations that need to succeed together in transactions. If anything fails, the transaction rolls back. ```typescript theme={null} await db.transaction(async (tx) => { const company = await tx.insert(companies).values({ id: data.id, name: data.name }).returning() for (const round of data.funding_rounds) { await tx.insert(fundingRounds).values({ company_id: company.id, round_name: round.name, amount: round.amount, }) } // If any insert fails, everything rolls back }) ``` **Useful error messages** When something fails, include context: what failed, why it failed, what happens next. Not "An error occurred." Instead: "Failed to sync company Acme Inc from PitchBook: Rate limit exceeded (429). Will retry in 60 seconds." ## Working with LLM Providers If you're building features that use LLMs: **Use an AI gateway for failover**. LLM providers have frequent outages and rate limits. Services like [Vercel's AI SDK](https://ai-sdk.dev/) or [LiteLLM](https://www.litellm.ai/) provide automatic failover between providers. **Handle streaming responses**. When building UIs, you'll want LLM APIs to stream tokens incrementally. **Set timeouts**. LLM requests can take 30+ seconds. Set 60-90 second timeouts so slow requests don't block indefinitely. ## Authentication Patterns **API tokens** Most vendors provide REST APIs authenticated with bearer tokens. Store tokens securely: * Use environment variables, not hardcoded in code * Use a secrets manager for production * Never commit tokens to git * Rotate periodically **User authentication for internal tools** For internal dashboards, prefer OAuth (Google, Microsoft) over managing separate passwords. For service-to-service authentication, use API tokens or service accounts. Don't use user credentials for automated processes. **Orchestrating data imports** For scheduled data imports, use orchestration tools: * **[Dagster](https://dagster.io/)**: Data orchestration, can schedule imports and transformations * **[Airflow](https://airflow.apache.org/)**: Workflow orchestration, similar to Dagster ## The Bottom Line Much of the technical work at VC funds is glue code. Validate everything with Zod or Pydantic. Respect rate limits with exponential backoff. Use transactions to avoid partial failures. For LLMs, use an AI gateway for failover. # Knowledge Graphs Source: https://buildingfor.vc/guide/part-3-technical-foundations/knowledge-graphs When relationship-heavy queries matter for VC - market mapping, competitor analysis, network effects - and how to approach implementation. ## Overview Companies don't exist in isolation. They have investors, employees, competitors, customers, suppliers, and potential acquirers. Founders move between companies. Investors co-invest together. Companies get acquired by larger companies in adjacent markets. Understanding these relationships is often as important as understanding the companies themselves. Traditional relational databases are optimized for storing structured data in tables. They work well for tracking companies, investments, and people as individual entities. But when you need to query relationships between entities (who invested in companies that competed with this portfolio company? which founders worked together at previous companies? what's the network of potential acquirers for this company?), relational databases become awkward. You end up writing complex joins across many tables, and performance degrades as you traverse multiple levels of relationships. Knowledge graphs are designed specifically for relationship-heavy data. They represent entities as nodes and relationships as edges. Queries that would require complex joins in a relational database become natural graph traversals. This makes them particularly powerful for certain VC use cases: market mapping, competitor analysis, finding similar companies, tracking networks of people and investors, and understanding acquisition patterns. But knowledge graphs and graph databases come with their own complexity. They're harder to work with than relational databases. Most engineering teams are more comfortable with SQL than with graph query languages like [Cypher](https://neo4j.com/docs/getting-started/cypher/) or [Gremlin](https://tinkerpop.apache.org/gremlin.html). Graph databases have different performance characteristics and require different thinking about data modeling. This chapter covers what knowledge graphs are, why they matter for VC, when to use them versus sticking with relational databases, and how to approach implementation without over-complicating your infrastructure. ## What Are Knowledge Graphs? A knowledge graph represents information as a network of entities (nodes) and their relationships (edges). Instead of organizing data into tables with rows and columns, you organize it as a graph where each node represents something (a company, a person, an investor) and each edge represents a relationship between two nodes (works at, invested in, acquired by). **Nodes** are entities. In a VC knowledge graph, nodes might represent companies, people, investors, funding rounds, or products. Each node can have properties. A company node might have properties like name, founding date, sector, and description. A person node might have name, title, and LinkedIn URL. **Edges** are relationships between nodes. These can be directional (Alice works at Acme Corp) or bidirectional (Company A competes with Company B). Edges can also have properties. An "invested in" edge might have properties like investment date, amount, and ownership percentage. A "worked with" edge between two people might have properties indicating when and where they worked together. ```mermaid theme={null} flowchart LR SEQ["Sequoia Capital
Type: Investor
AUM: $85B"] STR["Stripe
Type: Company
Sector: Fintech"] AUC["Auctomatic
Type: Company
Acquired: 2008"] PC["Patrick Collison
Type: Person
Role: CEO"] JC["John Collison
Type: Person
Role: President"] SEQ -->|"invested in
$2M Seed
2011"| STR SEQ -->|"invested in
$20M Series B
2012"| STR PC -->|"founded
2010"| STR JC -->|"founded
2010"| STR PC -->|"worked at
2007-2008"| AUC JC -->|"worked at
2007-2008"| AUC PC ---|"siblings"| JC ``` This is a trivial example, but it illustrates how knowledge graphs map entities and relationships. Nodes have properties (company sector, person role, investor AUM) and edges have properties (investment amount, dates). Graph queries can traverse these relationships: "find all companies founded by people who previously worked together" returns Stripe by following the "worked at" edges from the Collisons to Auctomatic. A real knowledge graph would have thousands of nodes and millions of edges, but the structure is the same. **Graph databases** are databases optimized for storing and querying graph data. They make it efficient to traverse relationships. Instead of joining tables, you follow edges. Queries like "find all companies within 3 degrees of separation from this investor that operate in the same sector" are natural in graph databases but painful in SQL. **The difference from relational databases**: In a relational database, you might have a companies table, a people table, an investments table that links companies to investors, and an employment table that links people to companies. To answer "which companies did investors who funded my competitor also fund?" requires joining investments table to itself through companies, then joining again to get company details. In a graph database, you start at your competitor node, follow "invested in" edges to investor nodes, follow their other "invested in" edges to company nodes, and you're done. ## Why Knowledge Graphs Matter for VC Venture capital is fundamentally about networks and relationships. Understanding these relationships helps with deal flow, diligence, portfolio support, and exits. **Market mapping**: When you're researching a sector, you don't just want a list of companies. You want to understand how they relate to each other. Which companies compete directly? Which are upstream suppliers or downstream customers? Which companies were founded by people who worked together previously? A knowledge graph makes it natural to visualize and query these relationships, turning a flat list of companies into a network that shows market structure. **Competitor analysis**: For portfolio companies or potential investments, understanding the competitive landscape means knowing not just who the competitors are, but how they're connected. Do they share investors? Do they target the same customers? Have they hired from each other? These relationship patterns reveal competitive dynamics that a flat database misses. **Company similarity**: Finding similar companies is valuable for sourcing (if we liked this investment, what else is similar?) and benchmarking (how does our portfolio company compare to similar companies?). Similarity based on relationships (same investors, same hiring sources, same partners) often reveals more than similarity based on keywords or sector tags. **[EQT's CompanyKG](https://github.com/llcresearch/CompanyKG2)** (1.17M companies, 51M edges representing 15 relationship types) focuses heavily on this, quantifying similarity through network structure and company description embeddings. They benchmarked 11 different methods, validating that relationship data is valuable enough to build serious infrastructure around. **Network analysis, M\&A patterns, investment patterns**: Tracking who knows whom, which investors co-invest, which companies acquire others, which founders worked together. Knowledge graphs make these network queries natural. Understanding acquisition paths, warm intro paths, and investor syndicate behavior all emerge from relationship data. **Identifying tech hubs**: Knowledge graphs reveal entrepreneurial concentrations. Which companies are "founder factories" (PayPal, Google, Stripe)? Which universities produce the most founders in your sectors? A graph query can show that 15 AI infrastructure companies were founded by people who worked at Meta's infrastructure team, suggesting a network worth watching. ## GraphRAG: Knowledge Graphs for AI Systems An emerging application of knowledge graphs in VC is GraphRAG (Graph Retrieval-Augmented Generation). This combines knowledge graphs with large language models to improve how AI systems retrieve and reason about your data. **What is GraphRAG?**: Traditional RAG systems retrieve relevant documents or text chunks based on similarity to a query, then feed those chunks to an LLM to generate an answer. GraphRAG enhances this by using a knowledge graph to understand relationships between entities. Instead of just retrieving similar text, it can traverse the graph to find related information that might not be textually similar but is semantically connected. **Why it matters for VC**: When a GP asks "which companies in our portfolio are working on AI infrastructure?", a traditional RAG system might retrieve documents mentioning "AI infrastructure." But GraphRAG can understand that Company A's infrastructure product is used by Company B (even if that relationship isn't explicitly stated in searchable text), or that the founders of Company C previously worked at Company D which built similar technology. The knowledge graph captures these relationships, making retrieval smarter. **Use cases**: * Answering questions about portfolio company relationships ("which of our portfolio companies could partner with each other?") * Understanding competitive landscapes ("show me companies that compete with our portfolio companies and their relationships to potential acquirers") * Connecting research insights ("what companies are working in adjacent spaces to this thesis we developed?") * Due diligence queries ("what are all the connections between this company and our existing network?") **How it works**: You build a knowledge graph of your companies, people, investors, and their relationships. When someone asks a question, the system: 1. Identifies relevant entities in the question 2. Traverses the knowledge graph to find related entities and relationships 3. Retrieves documents/data associated with those entities 4. Provides both the graph structure and the documents to the LLM 5. The LLM generates an answer informed by both textual content and relationship structure **Implementation considerations**: GraphRAG requires both a knowledge graph infrastructure and an LLM integration. You need clean entity resolution (covered in [Entity Resolution](/guide/part-3-technical-foundations/entity-resolution)) so the graph accurately represents reality. You need good data quality so the LLM doesn't hallucinate based on incorrect relationships. And you need to decide whether GraphRAG's benefits justify the additional complexity over simpler RAG approaches. For most funds, this is still emerging territory. Traditional RAG systems work well for searching documents and research. GraphRAG becomes valuable when relationship-based reasoning is critical to answering questions. If you're already building a knowledge graph for market mapping or competitor analysis, adding GraphRAG capabilities might be a natural extension. But don't build a knowledge graph just to enable GraphRAG unless you have validated that relationship-based retrieval solves problems traditional search doesn't. ## When to Use Knowledge Graphs vs. Relational Databases Knowledge graphs are powerful, but they're not always the right choice. The decision depends on your queries, your team's expertise, and your infrastructure maturity. **Use knowledge graphs when:** * Your core queries involve traversing multiple levels of relationships (find companies 2-3 hops away through investor networks) * You need to visualize network structure (market maps showing company relationships) * Relationship patterns are as important as entity attributes (co-investment networks, competitive clusters) * You're building recommendation systems based on graph similarity (companies similar to this one based on their network position) **Stick with relational databases when:** * Most queries are about entity attributes, not relationships (show me all Series A companies in fintech) * Your team is much more comfortable with SQL than graph query languages * You're still figuring out your data model and need flexibility to change it * Performance for your use cases is fine with relational databases **The hybrid approach (recommended):** For most VC funds, the right answer is to model everything in a relational database first, then build knowledge graph views or projections on top when you need relationship-heavy queries. **Why this works:** * Your team already knows SQL. Teaching everyone Cypher or Gremlin adds complexity. * Relational databases are well-understood, mature, and have great tooling. * Most of your queries are probably fine in SQL. It's only specific relationship queries that benefit from graphs. * You can start simple and add graph capabilities later when you validate the need. **How to implement this:** * Model your core data (companies, people, investors, deals) in Postgres or similar relational database * Ensure you capture relationships (investments, employment, board seats, partnerships) * For basic relationship queries, use SQL joins (they work fine for 1-2 levels of relationships) * When you need complex graph queries, either: * Build a materialized view in a graph database (periodically sync data from the relational database to the graph database) * Use graph algorithms directly on your relational data (libraries like [NetworkX](https://networkx.org/) in Python can build graphs from relational queries) * Query your relational database and build graph structures in your application layer when needed This gives you the benefits of graph thinking without committing to graph databases before you need them. ## Implementation Considerations If you do decide to use a graph database (either as your primary store or as a secondary system), here are key considerations. **Graph database options:** * **[Neo4j](https://neo4j.com/)**: Most popular graph database, mature ecosystem, Cypher query language. Good developer experience and tooling. * **[Amazon Neptune](https://aws.amazon.com/neptune/)**: Managed graph database on AWS, supports both Gremlin and SPARQL. Good if you're already on AWS. * **[Memgraph](https://memgraph.com/)**: In-memory graph database startup with focus on streaming data and real-time analytics. Compatible with Neo4j's Cypher query language. * **[FalkorDB](https://www.falkordb.com/)**: Graph database startup built on Redis, emphasizing low latency and high throughput. Also uses Cypher. Good for real-time applications. Note that Memgraph and FalkorDB are newer startups, whereas Neo4j and Neptune are more established with larger ecosystems and enterprise support. **Data modeling**: Graph data modeling is different from relational modeling. You need to think about what should be nodes vs. properties (is a funding round a node or a property of an investment edge?), how to model time-varying relationships (person worked at company from date X to date Y), and how to handle different relationship types with different semantics. **Query performance**: Graph databases are optimized for traversals, but performance depends heavily on your graph structure and query patterns. Highly connected nodes (like major investors) can become bottlenecks. You may need to denormalize data or add specialized indexes. **Data freshness**: If you're maintaining both a relational database and a graph database, you need to keep them in sync. This usually means ETL processes that run periodically (nightly or weekly). Real-time sync is possible but adds significant complexity. **Team expertise**: Graph databases require different mental models than relational databases. Your team needs to learn graph query languages, understand graph algorithms, and think differently about data modeling. This is an investment. Make sure it's worth it for your use cases. ## The Bottom Line Knowledge graphs are powerful for relationship-heavy queries: market mapping, competitor analysis, network effects, and similarity calculations. They represent how companies, people, and investors connect, which is often as important as the entities themselves. But don't start with a graph database. Most VC funds should model their data in relational databases first. Relational databases are well-understood, your team knows SQL, and they work fine for most queries. Only when you have validated use cases for complex relationship queries should you add graph capabilities. The hybrid approach works best: model everything in Postgres or similar, then build knowledge graph projections on top when you need them. This could mean materializing data into Neo4j for specific analyses, using graph libraries on top of your relational data, or building graph views in your application layer. EQT's CompanyKG demonstrates the value of knowledge graphs for VC at scale: 1.17 million companies, 51 million relationships, powering market mapping, competitor analysis, and company similarity. But they built this as a research project on top of their existing Motherbrain platform, not as a replacement for it. If relationship queries become core to your competitive advantage, invest in proper graph infrastructure. Until then, model your relationships clearly in relational databases and build graph capabilities as needed. In the next chapter, we'll cover integrations and APIs: common patterns for connecting VC tools, validating API responses, handling webhooks and rate limits. # Security and Compliance Source: https://buildingfor.vc/guide/part-3-technical-foundations/security-and-compliance Handling sensitive data, basic security principles, and working with compliance teams - what you need to know as a VC technologist. ## Overview VC funds handle sensitive data. Information about LPs (limited partners) who invest in your fund. Material non-public information (MNPI) about companies you're evaluating or have invested in. Portfolio company financials, cap tables, and metrics. Internal investment memos with proprietary analysis. All of this needs to be protected both for regulatory reasons and to maintain trust with LPs, portfolio companies, and the founders you work with. This chapter isn't a comprehensive guide to SEC regulations or compliance frameworks. Those require specialized legal and compliance expertise that goes beyond technical implementation. Instead, this chapter covers what you need to know as someone building technology for a VC fund: what sensitive data exists, why it matters, basic security principles you should follow, and when to involve compliance and legal teams. The bottom line: take security seriously, use proven tools rather than building your own, and work closely with your fund's compliance officers and lawyers on anything related to regulatory requirements. ## Sensitive Data and Why It Matters VC funds handle several types of sensitive data. Understanding what's sensitive and why it matters helps you make the right decisions about access controls, encryption, and retention policies. **LP (Limited Partner) PII** Your LPs are individuals and institutions who have invested money in your fund. You have their personal information: names, addresses, social security numbers or tax IDs, bank account details, investment amounts. This is personally identifiable information (PII) that must be protected for both regulatory and trust reasons. Private fund advisers managing \$150M or more in AUM must register as RIAs with the SEC, which comes with specific compliance obligations around data handling. Funds below this threshold can often qualify as Exempt Reporting Advisers (ERAs) with lighter requirements. But regardless of registration status, LPs invest millions of dollars in your fund and need to trust you'll handle their information properly. Data breaches with LP PII can result in lawsuits and damaged LP relationships. If you're building or buying fund operations software or LP portals, this data needs strong access controls. Most funds restrict LP data access to fund operations staff and partners. **Material Non-Public Information (MNPI)** When you're evaluating companies or working with portfolio companies, you learn things that aren't public: upcoming fundraising rounds, acquisition discussions, financial performance, product roadmaps, major hires or departures. This is MNPI, and mishandling it has legal consequences. SEC enforcement actions for improper handling of MNPI aren't theoretical - they happen, and they're expensive. For engineers, the practical implication is that any system touching company data (CRM, data warehouse, research platforms) needs proper access controls. You can't have this data accessible to everyone at the fund, and you definitely can't expose it publicly or to unauthorized parties. **Portfolio Company Data** Your portfolio companies share sensitive information with you: detailed financials, metrics, cap tables, customer lists, strategic plans. They trust you to keep this confidential. Leaking portfolio company data damages your reputation and makes it harder to work with companies in the future. Other founders hear about it, and your ability to build relationships suffers. When you're building portfolio dashboards or analytics systems, consider who should have access. Usually partners and investment team members need access, but not everyone at the fund. **Internal Investment Analysis** Your investment memos, due diligence notes, IC (investment committee) discussions, and proprietary research represent your fund's intellectual property. Competitors would love to know what you're looking at, what your thesis is, and how you evaluate companies. While this isn't regulated the same way as MNPI or LP PII, it's still sensitive and needs protection. When building or buying research platforms or CRM systems, think about access controls and who can see what information. **The bottom line**: You're handling data that has regulatory requirements (LP PII, MNPI), relationship consequences (LP trust, portfolio company trust), and legal liability (SEC enforcement, lawsuits). The specific regulations are complex and require legal expertise to interpret correctly. Work with compliance officers and lawyers who understand what applies to your fund's specific situation. ## Basic Security Principles **Use proven tools**: Don't write your own authentication, encryption, or access control. Use OAuth providers ([Google](https://support.google.com/cloud/answer/15549257?hl=en), [Okta](https://www.okta.com/)), authentication services ([Auth0](https://auth0.com/), [Clerk](https://clerk.com/)), and standard encryption libraries. Security is hard - use tools built by experts. **Encrypt data**: Use HTTPS everywhere (data in transit). Enable database encryption at rest (checkbox option in managed services). For LP PII or MNPI, consider application-level encryption. **Implement access controls**: Not everyone needs access to everything. Use role-based access control (RBAC): define roles (partner, associate, operations) and permissions (read portfolio data, edit CRM, view LP info). Use audit logs to track who accessed what. If you're building a multi-tenant application (for portfolio companies or LPs), implement Row-Level Security (RLS). **Principle of least privilege**: Give people and services only the minimum access needed. This limits damage if credentials are compromised. **Keep data only as long as needed**: Define retention policies with your compliance team. RIAs typically must keep some data 5-7 years. Implement automatic deletion for other data. ## Working with Your Compliance and Legal Teams As an engineer or data person building systems for a VC fund, you're not expected to be a compliance expert. But you need to work with people who are. **Before building systems that handle sensitive data, ask:** * What regulatory requirements apply to this data? (LP information has different requirements than CRM data about companies you're researching) * Who should have access to this data? (Define roles and permissions upfront) * How long should we keep this data? (Set retention policies early) * Do we need audit trails? (Usually yes for anything involving LP data or MNPI) * Are there specific security controls required? (Encryption, MFA, etc.) Your compliance officer or external compliance consultant can answer these questions. They understand RIA regulations, SEC requirements, and what your fund's specific obligations are. **When launching new tools, get compliance review**: Before you launch a new LP portal, portfolio dashboard, or research platform, have compliance review it. They'll check whether access controls are appropriate, whether data handling meets requirements, and whether you're missing anything important. This isn't adversarial. Compliance teams want you to build good tools. They just need to make sure those tools don't create regulatory or legal problems. **Document your security and access controls**: Write down who has access to what systems, what data is stored where, what encryption you use, and how long you retain data. This documentation is useful for audits, for onboarding new team members, and for compliance reviews. ## Practical Implementation Guidance Here's how these principles translate to the systems you're actually building. **For CRM systems**: * Use SSO (single sign-on) with your company's identity provider so you can centrally manage access * Set up role-based permissions so not everyone can see all data * Enable audit logging if the CRM supports it **For data warehouses** (Postgres, Snowflake, BigQuery): * Enable encryption at rest (checkbox in most managed services) * Use separate credentials for different services (your research platform should have different database credentials than your portfolio dashboard) * Limit access to production data - most people should use development/staging environments * Set up query logging so you can audit who accessed what data **For internal tools and dashboards**: * Require authentication: don't build unauthenticated internal tools that anyone with the URL can access * Use HTTPS everywhere * Use an established auth framework ([NextAuth](https://next-auth.js.org/), [Clerk](https://clerk.com/)) rather than rolling your own * Consider IP allowlisting for particularly sensitive tools (only accessible from office network or VPN) * Implement session timeouts so inactive sessions don't stay logged in forever **For API integrations with external vendors**: * Store API keys securely (environment variables or secrets manager, never in code) * Use separate API keys for production vs. development * Rotate API keys periodically (every 6-12 months) * Monitor API usage for anomalies (sudden spike in requests might indicate compromised credentials) **For data sharing with portfolio companies**: * Use authenticated, access-controlled systems rather than emailing spreadsheets * Use per-company access controls (Company A can only see their own data, not Company B's) * Track what data was shared with whom * Remove portfolio company access when projects finish ## If You Want to Go Above and Beyond The security principles covered above are sufficient for most small to mid-size VC funds. But if you want to implement additional security measures, or if your compliance team, LPs, or fund size require them, here are areas to consider: **SOC 2 compliance**: SOC 2 certification demonstrates that your systems meet specific security, availability, and confidentiality standards. This is common for companies selling software to enterprises, and some larger funds pursue it (especially if LPs explicitly require it). For most small to mid-size funds building internal tools, it's not necessary, but it can provide additional assurance to LPs and portfolio companies about your data handling practices. **Penetration testing**: Regular penetration testing by security professionals can identify vulnerabilities in your systems before attackers do. For most funds, following security best practices is sufficient, but pentest results can provide additional confidence and are sometimes required by compliance teams or LPs at larger funds. **Data Loss Prevention (DLP)**: Enterprise DLP systems monitor file transfers, emails, and data movement to prevent sensitive information from leaving your organization. For most funds, basic access controls and encryption are sufficient, but DLP can add an extra layer of protection if you're handling particularly sensitive data or operating at significant scale. **Advanced threat detection**: Enterprise-grade SIEM (security information and event management) systems aggregate logs from all your systems and use machine learning to detect potential security incidents. This is typically overkill for small to mid-size funds, but becomes more relevant at larger funds with sophisticated IT infrastructure. Ask your compliance team what's actually required vs. what's nice to have for your fund's specific situation. For most smaller funds, the basic security principles covered earlier are sufficient. These advanced measures become more relevant as you scale. ## The Bottom Line Security and compliance are important for VC funds. You handle sensitive data about LPs, portfolio companies, and your investment strategy. Mishandling this data has regulatory, legal, and reputational consequences. As someone building systems for a VC fund, follow these principles: * Use proven tools for authentication, encryption, and access control, don't build your own * Encrypt data at rest and in transit * Implement role-based access controls and the principle of least privilege * Keep audit logs of who accessed what data * Define data retention policies and delete data when it's no longer needed * Work with your compliance and legal teams on regulatory requirements Most importantly, recognize what requires specialized expertise. You don't need to become an expert in SEC regulations or RIA compliance. You need to build systems that make it possible to meet those requirements, and you need to work with people who understand the legal side. Get compliance review before launching tools that handle sensitive data. Document your security controls. Use managed services that handle security fundamentals for you. And when in doubt, ask your compliance team. In the next chapter, we'll cover emerging trends in VC technology: MCP, AI agents, and what's worth paying attention to versus what's hype. # Quick Reference Source: https://buildingfor.vc/guide/quickstart Find what you need quickly based on your role, experience level, or specific technical challenge. ## Reading Paths Different readers will find different parts of this guide most valuable. Choose your path: ### New to VC Funds If you just joined a VC fund or are new to the venture capital industry: Read all four chapters sequentially to understand how VC works, analyze your specific fund, avoid common mistakes, and build the right data team. Use [Understanding Your VC Fund](/guide/part-1-understanding-vc/understanding-your-vc-fund) to determine your fund's stage, volume, and strategy - this shapes everything you'll build. Apply lessons from [Common Mistakes](/guide/part-1-understanding-vc/common-mistakes) to avoid jumping in too early and building the wrong thing. ### Experienced with VC If you're already familiar with how venture capital works: Skip to the seven most common technical mistakes at VC funds. Jump to Part 3 for deep dives on data modeling, integrations, and architecture. ### Building Specific Features Looking for guidance on a specific technical challenge? * **[Data Modeling](/guide/part-3-technical-foundations/data-modeling)** - How to structure companies vs. deals * **[Data Providers](/guide/part-3-technical-foundations/data-providers)** - Which vendors to use and how * **[Data Quality](/guide/part-3-technical-foundations/data-quality)** - Validation, trust levels, and ground truth * **[Data Warehousing](/guide/part-3-technical-foundations/data-warehousing)** - When you need it and how to set it up * **[Integrations and APIs](/guide/part-3-technical-foundations/integrations-and-apis)** - Webhooks, rate limiting, validation * **[Entity Resolution](/guide/part-3-technical-foundations/entity-resolution)** - Matching companies across data sources * **[Data Providers](/guide/part-3-technical-foundations/data-providers)** - API vs file delivery, authentication * **[The VC Tech Stack](/guide/part-2-tech-stack/introduction)** - Complete survey of what tools matter * **[CRM & Deal Flow](/guide/part-2-tech-stack/crm-and-deal-flow)** - Choosing and implementing * **[Research Platforms](/guide/part-2-tech-stack/research-platforms)** - Building your thesis engine * **[Portfolio Support](/guide/part-2-tech-stack/portfolio-support)** - Different approaches and strategies * **[Security & Compliance](/guide/part-3-technical-foundations/security-and-compliance)** - What matters, what doesn't * **[Common Mistakes](/guide/part-1-understanding-vc/common-mistakes)** - Understanding confidentiality requirements (see [mistake #5](/guide/part-1-understanding-vc/common-mistakes#mistake-%235%3A-not-understanding-data-sensitivity-and-compliance)) ## Common Questions This depends on your fund's stage, resources, and specific needs. See: * [Common Mistakes, #2](/guide/part-1-understanding-vc/common-mistakes#mistake-%232%3A-not-aligning-on-buy-vs-build-strategy) * Part 2: The VC Tech Stack for category-by-category guidance [Choosing Your Stack](/guide/part-3-technical-foundations/choosing-your-stack) covers technology choices with VC-specific reasoning. The short answer: use boring, proven technology that AI coding tools understand well (TypeScript/Python, Next.js, Postgres). [Data Modeling](/guide/part-3-technical-foundations/data-modeling) covers the critical distinction between companies and deals, plus all the entity types you need to track. This is one of the most common areas where developers go wrong. [Data Providers](/guide/part-3-technical-foundations/data-providers) covers data providers by category (company data, people data, signals), plus how to evaluate vendors and work with their APIs. [Security and Compliance](/guide/part-3-technical-foundations/security-and-compliance) covers what's actually sensitive, what audit requirements matter, and how to balance security with productivity. Also see [Common Mistakes, #5](/guide/part-1-understanding-vc/common-mistakes#mistake-%235%3A-not-understanding-data-sensitivity-and-compliance).