🎁 Zanim zaaplikujesz, przećwicz tę rozmowę. Załóż darmowe konto WorkMundi i odbierz Trening rozmowy kwalifikacyjnej w HelpsYouSpeak — bez opłat, bez karty. Chcę swój trening →
InPost Group is an innovative European out of home deliveries company, revolutionizing the way parcels are delivered to customers. With operations across several countries, our network of intelligent lockers provides customers with a fast, convenient, and secure delivery option. InPost Group is a publicly traded company, with a market capitalization of about $5 billion as of March 2023. With over 10,000 people worldwide, InPost Group is one of the largest out of home delivery providers in Europe, committed to providing sustainable and efficient delivery solutions to meet the evolving needs of customers in today's rapidly changing landscape. We are a team that builds the best FMCG e-commerce in Poland. We already have over 300,000 products in our offer, we work with over 300 sellers , and our mobile application has already been downloaded by almost 1.5 million users in Poland . Due to the fact that our product consists of many elements that must work together efficiently and be managed effectively, we are looking for experienced consumer-focused (Web / Mobile App products) engineering leaders to join us in that journey - heavily influence our future platform build, improve processes and help us deliver best customer experience in the market across all types of devices. Why this role exists Von Halsky is InPost's conversational AI shopping assistant, live in production and serving a growing share of our customers. What decides whether it wins is not the model but whether we can tell, at release cadence, that a change made conversations better. That is the Evaluations Platform, and we are hiring the engineer who takes it to the next level. What you will own The LLM-judge pipeline. Our release-gating judge over real and golden conversations, calibrated well enough that people act on its verdicts instead of arguing about it. Eval datasets and the golden set. Real coverage across intents and categories, including the Polish-language coverage generic benchmarks do not give us. Root-cause analysis on real conversations. Making our conversation-mining stack diagnostic rather than descriptive, on a stable issue taxonomy. The "sus" detector. Abusive, adversarial and anomalous sessions, next to our guardrails and red-team work. Evals-driven development. The eval comes before the feature, and writing it is as cheap as writing the code. The interface to product and business. Vague asks in, measurable quality definitions out, and results stakeholders can act on. What "leadership aspirations" means here concretely This is a technical lead role, not a people-management role, and you will not carry line-management duties on day one. You will set and defend the technical direction for the platform, act as reviewer of record for the area, mentor other engineers, scope work with our PM and EM, present results to stakeholders, and hold the line against ad-hoc requests crowding out platform work. Engineering management later, or a Staff-level hands-on track, are both paths we will build with you. Either way we need someone accountable for an area rather than for a ticket. How we define success in this role Judge pass rate is calibrated against human labels and is a number people trust and cite. A stable, versioned issue taxonomy is live and week-over-week trends are comparable. Every release is gated by an eval run the team can reproduce. The golden set has documented coverage and named blind spots. At least two engineers besides you can operate and extend the platform. What we are looking for Required 5+ years building and running production software, with 2+ years on LLM-based systems that real users hit. Strong engineering fundamentals , plus the habits that go with production ownership: testing, CI/CD, containers, observability, and working in cloud. We work primarily in Python. Demonstrable experience evaluating generative systems, not only building them: LLM-as-judge, human-label calibration, inter-annotator agreement, regression suites, offline-versus-online divergence. You should have opinions about what makes an eval worthless. AI engineering fundamentals. Prompt and context engineering as a discipline (context-window budgeting, structured outputs, failure-mode taxonomies); agentic primitives in production (tool use, multi-turn state, MCP , agent-to-agent integration patterns); and eval and LLM-observability tooling (LangFuse, Braintrust, Weave or equivalent, including things you built yourself). Comfort with data at scale : SQL, working with a lake or warehouse, and building a metric someone else can reproduce. Fluency with AI-assisted development tooling (Claude Code, Cursor, Copilot). We use it daily and expect it. Ability to make a technical argument to a non-technical audience and be understood. English B2 and Polish . Our users converse in Polish and you will read their conversations; judging quality you cannot read is not possible. Nice to have Harness and loop engineering. Building the scaffolding around models rather than only calling them: agent loops, retries and fallbacks, tool-call orchestration, deterministic replay, and the plumbing that makes a non-deterministic system testable. Auto-improving systems. Closing the loop from production signal back into the product: mining failures into cases, using eval results to drive prompt, retrieval and routing changes, and automating the parts of that cycle that people do by hand today. Adversarial robustness, jailbreak testing, red-teaming, or abuse and fraud detection. E-commerce, search or recommendation domain experience. What we offer A product that has already been released to millions of users, with a real business case, not a lab pilot. Direct access to frontier models at committed capacity across multiple providers, plus an open-source track we run ourselves. A quality mandate with executive attention. Ownership of a platform that is greenfield in practice inside a company with production traffic, which is the rarest combination in this market. Hybrid working from Warsaw or Kraków, in a team growing fast enough that early hires shape how it works. Fulfilling careers with a range of benefits for people and investing in providing training opportunities for their development. You will feel a part of the InPost community that makes an impact on sustainability, convenient deliveries, and the circular economy every day. Excellent working environment and flexible hours We offer B2B type of contract
Here's how to pick the right one and stand out in your application.
144.883Jobs
31.687IN
81%EN
That number is real. WorkMundi's database shows 144,883 open engineer roles across the world. India has the most with 31,687 jobs, followed by the United States with 30,084. If you just finished reading one job ad and felt paralyzed by choice, you're not alone—but this scale is actually an advantage. It means you can afford to be selective.
Start by geography and language. The majority of engineer ads—117,837 of them—have the job posting text written in English. Use that as one filter, but remember: the ad text language tells you nothing about whether the role actually requires you to speak English day-to-day. Read the job description carefully. Then check which countries have the volume you're targeting. Singapore, Poland, and Australia round out the top five after India and the US.
Next, learn who's hiring. Accenture has posted 2,801 engineer roles. andurilindustries, speechify, and jobgether are also actively recruiting. If you're applying to one of these names, research their hiring patterns and interview style before you apply. That homework pays off.
When you interview, expect the question every engineer hears: 'Tell me about a time you had to debug a problem that wasn't in your job description.' Have a specific story ready—not a general one. Name the tools, the deadline pressure, and what you learned. Hiring managers listen for whether you see problem-solving as part of the role itself, not a favour.