🎁 Before you apply, rehearse this interview. Create your free WorkMundi account and get an Interview Training on HelpsYouSpeak — no cost, no card. I want my training →
Job Title: Founding Engineer - Privacy & PII [Data Products] Location: On-site, Bengaluru Employment Type: Full-Time About the role Our entire product rests on one promise: raw enterprise data goes in, and what comes out is safe to sell nothing identifiable, nothing leaked, provably de-identified, with the audit trail to back the claim. You'll own that promise. You'll build the detection, entity-consistent pseudonymization, vault, and residual-risk gates that convert sensitive enterprise data into a trainable product buyers and regulators can trust. This is a hands-on founding role and the most trust-critical seat on the team. The difference between a sellable product and a liability is the judgment in this role. Responsibilities Own the de-identification pipeline end to end: sensitive-entity detection entity-consistent pseudonymization residual-risk validation clean exportBuild PII/PHI/secret detection across free text, structured fields, and layout, using and extending tools like Presidio plus custom recognizersDesign and operate the vault: one durable, namespaced surrogate per resolved entity, so the same person maps to the same pseudonym across every system (name, email, employee id, Slack/Jira/Git ids) without token driftBuild the "rewrite all surfaces" logic that applies surrogates consistently across the datasetDesign and enforce the residual-identifier release gate (regex + NER + secret scanners + sample review) and make explicit where it does and doesn't establish legal anonymizationEnforce the trust boundary: raw stays encrypted inside the customer environment; vault keys never leave; buyers receive only the cleaned productBuild fail-closed behavior (no raw pass-through when the vault or detection is unavailable) and full audit/lineage on every reveal and exportProduce the method cards / provenance records that document de-identification technique and residual-risk posture for buyersDesign detection and compliance profiles as swappable configuration so new verticals (medical/PHI later) extend the system rather than rewrite itPartner with the Data Products engineer at the handoff point, and with privacy counsel on what the product can honestly claim Must-have skills 6+ years engineering, including direct experience building PII detection, de-identification, tokenization, or privacy-aware data processing (this is the non-negotiable filter general data-engineering experience does not substitute)Strong Python; comfort with NLP/NER approaches for entity detection in unstructured textDeep, demonstrable understanding of the concepts that make or break this product: entity resolution before tokenization, token drift and why consistent surrogates matter, quasi-identifier re-identification risk, and the difference between pseudonymization and anonymizationExperience designing tokenization/pseudonymization systems with a durable mapping store (vault) and format/consistency guaranteesWorking knowledge of privacy regimes relevant to enterprise and health data (GDPR/EDPB framing, HIPAA Safe Harbor vs. Expert Determination) and how to engineer against them enough to know what can and can't be claimedAbility to balance aggressive de-identification against preserving the data's training utility (over-redaction destroys value; under-redaction destroys trust)Comfortable owning a trust-critical system and its quality bar in a small founding team Nice-to-have skills Presidio internals, custom recognizers, or comparable NER-based detection at scaleVault/tokenization backends (HashiCorp Vault Transform, Skyflow, Protecto, or equivalent) and format-preserving encryptionCloud DLP tooling deployed in-VPC (AWS Comprehend/Macie, Azure PII, Google DLP)Healthcare de-identification / HIPAA experience (directly relevant to our medical vertical)Differential privacy or synthetic-data familiarity for high-risk aggregate casesData lineage/audit tooling .
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.