AI integration that knows when to ask a person.
We connect language models to the systems you already run - documents, ERP, inbox, ticketing - for the jobs they do well: answering from your own files, pulling data out of documents, sorting incoming requests. Every feature ships with a test set, a human fallback and a visible cost per request.
- Assistants that answer from your documents and cite the source.
- Data pulled out of documents, with a person checking the doubtful ones.
- Anything that sends, books or pays waits for a person’s approval.
- EU-hosted or self-hosted models, personal data masked before any call.
Tested on your data, not on a demo
AI projects that disappoint tend to fail in the same place: a demo that worked on ten hand-picked examples meets a thousand real ones - crooked scans, emails in three languages, a policy replaced last spring. So before writing a prompt we collect a few hundred real examples from your data, agree what a correct answer is, and turn that into a test set that runs on every change.
Then we put the model where the work already happens - behind your API, inside the ERP screen, on the shared inbox. Models are good at reading, sorting and drafting; they are bad at arithmetic and at knowing what they do not know. The code around them covers that: citations, validation rules, confidence thresholds, permissions and an audit log.
We are not tied to a vendor: hosted models from OpenAI, Anthropic, Google or Mistral in EU regions, or open-weight Llama, Mistral or Qwen on your own hardware. Switching is a configuration change and one test run.
How it works
What AI can do for your team today
Not a robot that replaces anyone - three dull jobs it does well, with a person in charge. Shown at an equipment supplier.
-
A customer asks about delivery times. The assistant answers from your actual terms and shows the page it found the answer on - no made-up facts.
-
A supplier invoice arrives as a PDF. The AI pulls out the supplier, amounts and dates, and flags the one field it is not sure about.
-
The AI drafts a reply to an unhappy customer and suggests a credit note. Nothing goes out until someone on your team clicks approve.
How long does delivery to Vienna take?
Searching your documents…
Delivery to Vienna takes 3 working days. Orders placed before 12:00 ship the same day.
Source: Terms of delivery, page 22.1Germany: 4 working days.
2.2Austria: 3 working days. Orders before 12:00 ship the same day.
2.3Switzerland: 5 working days, customs included.
In practice
Illustrative scenarios - the kind of work we take on, not client case studies.
Service assistant for an equipment manufacturer
- The problem
- Technicians search manuals, warranty terms and old tickets on a shared drive from their phones. Finding the current revision usually ends with a call to the office.
- What we build
- An assistant over manuals and resolved tickets that respects folder permissions, cites document, page and revision, and says so when the documents do not cover a question.
- Python
- FastAPI
- PostgreSQL
- pgvector
- Azure OpenAI (EU)
- SvelteKit
Accounts-payable inbox for a distributor
- The problem
- Supplier invoices arrive as PDFs and scans in several languages. Two people retype them into the ERP, and mistakes surface at month-end.
- What we build
- Extraction into the ERP’s purchase-invoice fields with checks on totals, VAT, IBAN and the purchase order. Clean documents post automatically; the rest queue for review with the doubtful fields marked.
- TypeScript
- Node.js
- Tesseract OCR
- Mistral (EU)
- PostgreSQL
- ERP REST API
Ticket triage for a software vendor
- The problem
- Every support ticket lands in one queue, so an outage report can wait behind twenty password resets.
- What we build
- A classifier that tags product and urgency, routes to the right team and drafts a first reply from the knowledge base. Low-confidence tickets stay in the general queue; accuracy is checked weekly against a labelled sample.
- Python
- Llama (self-hosted)
- vLLM
- Redis
- Helpdesk API
- Grafana
What you get
- An evaluation set from your real data, with an agreed accuracy target per task
- Integration into your existing systems via API, webhooks or plugin
- Search over your documents that respects who may see what
- Guardrails: masking of personal data, allowed actions only, approval steps
- A review screen for low-confidence results that feeds the test set
- Monitoring of accuracy, speed and cost per request
- A setup you can move to another model or your own servers
Typical stack
- Python
- TypeScript
- FastAPI
- Node.js
- PostgreSQL + pgvector
- Qdrant
- OpenAI / Azure OpenAI
- Anthropic Claude
- Mistral
- Llama / Qwen (self-hosted)
- vLLM
- Ollama
- Langfuse
Questions people ask
Your question is not on the list? Ask us directly - we reply within two working days.
Ask a questionHow much does AI integration cost?
It depends on the number of tasks, the state of your data, the accuracy you need and how deep the integration goes. Running cost is mostly tokens per request times volume, which we measure on a sample of your data. We reply within two working days with a link to a 30-minute intro call, then send a written plan with a fixed-scope estimate.
How long does it take to build an AI chatbot for our company?
A first version that answers from your documents is the quick part. Permissions, review screens, integration and enough testing to trust it are most of the work. You get a timeline in the written plan, once we have seen the data and know where the answers need to appear.
Is our data safe, and will it be used to train the model?
The business APIs we use do not train on your requests by default, and we choose EU regions where they are offered. Personal data can be masked before any call, and access follows your existing permissions. If data must not leave your network, we run an open-weight model on your own servers.
How accurate is it, and what happens when the AI is wrong?
It will be wrong sometimes, so we plan for it. Accuracy is measured on your examples, not on a vendor benchmark. Below an agreed confidence, results go to a person; citations and validation rules catch most of the rest. Every correction becomes a new test case.
Which model do you use - ChatGPT, Claude or open source?
Whichever passes your test set at the lowest cost and an acceptable speed. Often a small model sorts and a larger one handles the complex answers. The provider sits behind one interface, so moving to a self-hosted Llama or Qwen is a configuration change plus a test run.
Bring one real task and a pile of real examples.
Tell us which task eats your team’s time and what the data looks like. We reply within two working days with a link to a 30-minute intro call, then send a written plan and a fixed-scope estimate - including how we will measure that it works.
Discuss your use case