Skip to main content
Arcus

Approach

Measure the fetch before arguing about the answer.

Why Arcus starts at retrieval rather than at generation, where relevance judgements actually come from, the boundary it keeps, and the list of ways this premise could turn out to be wrong.

Write to [email protected]. Replies go out within five working days.

First principle

1The order of operations

Fix the input before tuning the thing that reads the input.

A retrieval augmented pipeline has a strict dependency. The generation step can only reason over the passages the retrieval step handed it. That makes the quality of the fetch an upper bound on the quality of everything after it, and it makes tuning the generation step before measuring the fetch a way of spending weeks improving how well a model reasons about the wrong documents.

This is not a controversial claim in information retrieval. It is close to the founding assumption of the field. It is nonetheless routinely ignored in practice, and the reason is not ignorance. It is that the generation step is easy to inspect and the retrieval step is not. Anyone can read an answer and form an opinion. Reading a ranked list of forty passage identifiers and forming an opinion about whether the right one is missing takes a labelled set, and on most teams nobody has ever sat down and assembled one.

So the work is not persuading anyone that retrieval matters. Most engineers running these systems already believe it. The work is making the measurement cheap enough that it happens on an ordinary week rather than after an incident.

The hard part

2Where relevance judgements come from

This is the expensive, unglamorous centre of the whole thing, and it is worth being honest about it.

Every measure on the front page needs a set of judgements underneath it. For a given question, which documents in the corpus actually answer it, and how well. Somebody has to decide that, and the somebody has to be a person who understands the domain.

Why it cannot simply be automated away

The obvious shortcut is to ask a language model to judge relevance. It is a real technique, it is used in published work, and it is genuinely useful for expanding a small labelled set. It is not a replacement for human judgement, for a reason that matters here more than usual. If the same family of models is used both to retrieve and to judge, the evaluation inherits the retriever's blind spots and reports back that everything is fine. An instrument built from the thing it is measuring is not an instrument.

Our position is that a small set of human judgements beats a large set of automatic ones, that automatic judgements are a reasonable way to widen a human labelled set once agreement between the two has been measured on a sample, and that the agreement measurement is not optional. That is why label agreement is on the shelf next to the measures rather than hidden in an appendix.

Pooling, and the documents nobody judged

A corpus of two hundred thousand documents cannot be judged exhaustively against three hundred questions. The standard answer, used at the Text Retrieval Conference since the beginning, is pooling. Run several different retrieval configurations, take the union of the top results from each, and judge only that pool. Everything outside the pool is treated as not relevant.

That treatment is a known approximation and it has a known bias. A genuinely relevant document that no configuration in the pool ever retrieved is silently scored as irrelevant, which flatters every configuration equally and hides the same blind spot in all of them. Any honest report has to say how deep the pool was and how many configurations contributed to it, because those two facts bound how much the numbers can be trusted.

What this means for a buyer

The labelling cost stays with you, and it is worth knowing that before you write. Arcus makes the judging pass fast, samples so that the reading goes where it changes a number, and tells you when your judges disagree. It cannot make the reading go away, and a pitch that suggested otherwise would be the tell that the pitch was written by someone who had never done it.

Boundary

3Why the product stops where it stops

The most useful thing a small company can publish is the list of things it has decided not to do.

The pressure on anything in this space is to expand. Measure retrieval, then measure answers, then store the documents so measuring is easier, then serve the retrieval itself since you are already holding the documents, then host a model because the customer asked. Each step is individually reasonable and the destination is a product with no shape.

The boundary here is a single sentence. It reads result lists and it computes numbers. Everything on the far side of that sentence is somebody else's job.

  • No generation, ever. Not summaries, not explanations of a score in prose, not a chat interface over the results. A number and a sorted list of the worst questions is the whole output.
  • No index and no corpus. Document text is never requested, and there is no field for it to sit in.
  • No hosted retrieval. The pipeline being measured stays yours. An evaluator that also serves the traffic it evaluates is marking its own homework.
  • No performance monitoring. Latency and cost are worth measuring and are already well served. Mixing them in here would let a fast pipeline look accurate.
  • No claim about model quality. If the passages were right and the answer was wrong, this says so and then goes quiet.

Put a limit in writing and it becomes something you can be held to. Leave it in the founders' heads and it is only a taste, which is the kind of thing that quietly rearranges itself the week a big enough contract depends on it rearranging.

Risk

4Ways this premise could be wrong

Five of them, published rather than kept quiet, because the fifth one is the one that worries us.

  1. Context windows keep growing until retrieval stops mattering. If a model can read the whole corpus, ranking is a cost optimisation rather than a correctness problem. We think cost and latency keep retrieval alive at any realistic corpus size, but that is a belief and not a measurement.
  2. Nobody will pay for the labelling. The measurement is only as good as the judgements, the judgements need domain experts, and domain experts are the most expensive people on the payroll. A tool that requires a week of their time may simply never be used, however good it is.
  3. The platforms absorb it. Retrieval evaluation is a natural feature for whoever already sells the index or the model. If it arrives bundled and adequate, a separate tool has to be considerably better than adequate.
  4. Open source already covers it. There are respectable open libraries that compute these measures. If the gap is genuinely only the harness and the report, that gap may be too small to build a company inside.
  5. Teams do not actually want to know. A measurement that says the pipeline shipped last quarter retrieves the right document for six questions in ten is unwelcome information. Unwelcome information has a poor adoption record.

None of those are resolved, and a page giving only the reasons something works is marketing rather than method. Each one is written here in the form that would make us change course if it turned out to be true.

If one of them looks more likely from where you sit, or there is a sixth we have missed, [email protected] reaches a person who wants to hear it.

Replies go out within five working days.

The company

5Register facts, verifiable at source

Everything below can be checked against a public register rather than taken from us.

Registered name
ARCUS AI PTY LTD
Company form
Proprietary company, registered in Australia
ACN
697 547 505
ABN
82 697 547 505
Goods and services tax
Registered
Home jurisdiction
New South Wales
Trades as
Arcus, which is a name used by the company above rather than a separate entity
Verification
Entity status, the ABN and the GST registration can be looked up at no charge through abr.business.gov.au. The ACN record itself is held by ASIC at asic.gov.au
Service of documents
For formal service, the address with legal effect is the registered office ASIC holds against ACN 697 547 505. Adding a second postal address to this page would carry none of that effect

No postal address is published anywhere on this site. Under the Corporations Act 2001, section 153, a company must set out its name and its ACN on public documents, and that is what the footer of every page does. Nothing in that Act obliges a company to put a registered office address on a website, and where an address carrying legal effect is genuinely needed, the register itself is the authority. This page points at the register instead of restating an address that would add nothing to it.