PROJECT 17 / 18AGENT SYSTEMS RESEARCHPYTHON

Private · Phase 0 research prototype

Junction.

The right document can change the right tool.

2 resource typespassive context + executable tools
Exact subsetssmall-pool reference selector
Phase 0controlled experimental plumbing
01 / IDEA02 / SYSTEM03 / PLAYGROUND04 / DECISIONS05 / SOURCE
01 / THE IDEA

A closer look.

Junction studies joint selection of passive context and executable tools under shared budgets. Its current reference implementation includes typed resources, lexical candidate generators, an exact small-set selector, independent baselines, controlled toy tasks, and versioned execution traces.

A document and a tool can be valuable together even when neither ranks first alone. A reimbursement policy, for example, can make one API correct and another inappropriate. Junction makes those cross-resource interactions part of the selection objective instead of retrieving each resource type independently.

01

Represent different resources honestly

Context resources carry text and token cost. Tools carry a description, schema token cost, estimated latency, and estimated monetary cost. Both compete for the same visible context budget.

02

Optimize the combination

The exact selector enumerates small candidate subsets, rejects budget violations, and adds unary utility to context–tool interaction scores. Stable tie breaking makes toy decisions reproducible.

03

Compare to independent routing

Fixed and adaptive split baselines rank contexts and tools independently and deliberately omit pairwise interactions. These are controlled comparisons, not evidence of real-agent gains by themselves.

04

Keep evaluation labels out of the view

BenchmarkTask.router_view omits gold annotations and verifier configuration. Trace records separately track candidate pools, exposed resources, token usage, tool calls, retrieval metrics, and verified outcomes.

02 / UNDER THE SURFACE

Compile a useful resource set

Two retrieval lanes meet at a shared budget and a cross-type objective.

DRAG TO PAN · SELECT A NODE · + / − TO ZOOM

Read the architecture as text
  1. Task + visible state — The router view includes objective, available resources, budget, and initial state. Gold resource annotations and verifier configuration stay outside this inference-time view.
  2. Context resources — ContextResource is immutable and validates nonempty IDs and nonnegative token cost. Source metadata distinguishes policy, code, memory, or another origin without changing its passive role.
  3. Tool resources — ToolResource carries description, schema, schema token cost, estimated latency, and estimated cost. Its token_cost property exposes schema size to the same budget accounting as context text.
  4. Context candidates — LexicalContextCandidateGenerator scores the query against context content and source, then orders candidates by descending score and stable resource ID before applying the limit.
  5. Tool candidates — The separate tool candidate generator operates over executable resource descriptions and schemas. Keeping this lane distinct supports independent retrieval measurements before joint selection.
  6. Unary utility — The selector accepts a score mapping keyed by resource ID. For each feasible subset it sums the utility of included resources; absent scores default to zero. The current interface does not itself learn a utility model.
  7. Cross-type interactions — For every chosen context and tool, the selector adds scores for both ordered pair keys. Positive compatibility can reward a correct policy/tool pairing; negative interactions can penalize a conflicting combination.
  8. Shared resource budget — Budget is a typed set of hard limits. The exact selection loop enforces shared visible tokens, tool count, and summed tool latency/cost estimates. Actual call-count accounting is represented separately in traces.
  9. Exact subset search — ExactJointSelector enumerates combinations of candidate resources, applies budget filters, and compares the resulting objective. Exponential enumeration is intentional for a small reference problem, not a scalable production retrieval algorithm.
  10. Exposed resource set — Selection separates chosen context IDs and tool IDs, records token use, and exposes unary, interaction, and total objective scores. This allows the benefit of coupling to be inspected directly.
  11. Independent baselines — The fixed baseline reserves a token fraction for each resource type and greedily ranks them separately. The adaptive version searches a predefined split set without access to interaction information.
  12. Controlled toy tasks — The coupling example contains a contractor policy, generic text, an expense tool, and a payroll tool. Its explicit positive and negative pair scores isolate a case where independent relevance misses the intended mixed pair.
  13. Task outcome verifier — verify_toy compares a returned result with the generated task’s verifier specification. This evaluates controlled toy outputs; it is not evidence that a deployed language-model agent completed the task.
  14. Versioned raw traces — TaskTraceRecord validates candidate/exposed set relationships, contiguous tool-call indexes, usage consistency, success/failure labels, and schema version. JSON serialization refuses nonfinite numerical values.
03 / INTERACTIVE STUDY

Make the pair matter

Select from a small illustrative pool of documents and tools. Adjust the shared token budget, tool slots, and coupling strength to see when the best combination changes.

CHANGE THE INPUTS

A small illustrative exact-selection problem. Resource costs and utilities are demonstration values, not learned scores or experimental results. Passing a budget check does not establish task success.

ILLUSTRATIVE MODELLIVE

04 / ENGINEERING CHOICES

Why it works this way.

01

Price tool schemas as context

A tool consumes context before it is ever called. Treating its schema tokens and a document’s content tokens as one visible budget prevents free tools from distorting a comparison.

02

Use an exact reference first

Enumerating every subset is expensive at scale, but gives a transparent reference optimum for tiny problems. The current phase validates the scoring and accounting machinery before a larger solver or costly model study.

03

Separate labels from available information

Gold resource sets and verifier configuration are available to evaluation, not router_view. Raw traces retain enough structure to distinguish retrieval coverage from execution outcome.

05 / OPEN THE SOURCE

Trace it back.

Implementation details, examples, and project documentation.

Exact selector

Feasibility filters, unary utility, cross-type interaction terms, and subset comparison.

Candidate generation

Separate lexical retrieval contracts and stable tie breaking.

Independent baselines

Fixed and adaptive resource splits without interaction scores.

Experimental boundary

Phase 0 scope and the open research question about equal-budget agent success.

Scope & limitations

  • The source repository is currently private. This public case study explains its architecture; source files are not accessible to visitors.
  • The current repository is a Phase 0 reference and toy-evaluation system. It does not establish that joint retrieval improves real language-model agent success.
  • Exact subset search scales exponentially. The selector uses provided scores and estimated resource costs; these are not a learned utility model or a guarantee of runtime usage.

Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.