PROJECT 09 / 18BROWSER / AI SECURITYJAVASCRIPT / REACT

Browser extension prototype

Sentiency.

Inspect suspicious instructions before they blend into a workflow.

Manifest V3extension platform
0.65default model threshold
3remediation modes
01 / IDEA02 / SYSTEM03 / PLAYGROUND04 / DECISIONS05 / SOURCE
01 / THE IDEA

A closer look.

A Chrome extension prototype that combines local text heuristics with Gemini classification to flag potential prompt injection. It connects hidden-DOM scanning, clipboard handling, image transcription, and conversation monitoring to visible review and remediation controls.

Instructions can arrive inside a webpage, copied text, an image, or a conversation. A browser-side inspection layer makes the source, suspected spans, and proposed handling visible at the point where a user encounters the content.

01

Several browser entry points

Hidden-page observations, text paste, copying, manual scans, images, and supported conversation views feed shared analysis functions.

02

Layered text classification

Local Unicode, keyword, imperative, and encoding signals are combined with structured Gemini output. The confirmation rule and severity thresholds are explicit in source.

03

Reviewable span handling

Character spans are merged, bounded, and augmented before display or removal. Page highlighting and paste insertion have separate policies.

04

Image-to-text path

The primary image flow transcribes first, then classifies the extracted text. A combined vision path handles specified fallbacks and suspected visual-only content.

02 / UNDER THE SURFACE

Observe → classify → review

Several browser event sources feed one shared pipeline. Model confidence and local suspicion are separate signals.

DRAG TO PAN · SELECT A NODE · + / − TO ZOOM

Read the architecture as text
  1. Hidden-page scanner — The DOM engine focuses on visually concealed text and hidden images with meaningful metadata. Debounced mutation handling discovers newly added content. It does not imply complete coverage of all page instructions. The snippet is an illustrative data flow.
  2. Clipboard / selection — Clipboard and manual-selection paths send text into the shared pipeline. Paste handling can delay insertion until the configured remediation path resolves the result. The snippet names the inspected pipeline entry point.
  3. Image transcription — The primary image path requests structured OCR output, orders text blocks, then runs the text classifier on the resulting transcript. It falls back to combined vision classification for specified OCR failures or visual-only risk.
  4. Conversation monitor — On supported LLM sites, the session monitor waits for text stability, analyzes assistant messages, and checks a recent conversation trajectory every third processed assistant turn.
  5. Local detectors — analyzeText runs Unicode, instruction-pattern, and encoding detectors. Instruction suspicion is the capped sum of keyword hits at 0.3 and imperative matches at 0.2 each.
  6. Obfuscation unwrap — If encoding findings exist, the pipeline computes a decoded form. It passes original and decoded text to the classifier instead of silently replacing the original span coordinate system.
  7. Gemini classifier — The pipeline calls Gemini when there is local suspicion, text length above 150, or an explicitly forced source such as paste, scan, or copy. Source text may therefore be sent to an external model API.
  8. Trajectory analysis — Trajectory analysis slices recent turns, builds a multi-turn prompt, and asks Gemini for a trajectory result. It is a separate path from character-span text classification.
  9. Confirmation gate — A text result is confirmed when the model flags an injection at or above the configured threshold, or the local instruction score reaches 0.8. A low model confidence does not override a strong local signal.
  10. Classification + severity — The model’s class and technique map into a taxonomy path. Severity uses confidence cutoffs of 0.6, 0.75, and 0.9, with a one-level bump when Unicode and encoding anomalies both occur.
  11. Span reconciliation — Model-provided character spans are merged and clamped, then augmented with bracketed blocks, encoding segments, and local instruction hits. These reconciled spans drive highlighting or removal.
  12. Threat record — Confirmed threats are logged locally and announced to the extension service worker. The content-page path also dispatches a custom event so the in-page interface can react.
  13. Page handling — For live DOM content, non-BLOCK modes visualize suspected spans or outline media rather than surgically removing text. BLOCK can replace content or image metadata. The snippet is an actual mode branch.
  14. Paste handling — Clipboard handling applies the selected mode. Surgical handling removes resolved spans; if that would remove all text, the current implementation inserts the original and highlights it instead.
  15. Activity and review UI — The extension provides an activity side panel, settings, and in-page review components. The interface exposes the suspected source, taxonomy, confidence, and spans to the user. The snippet describes the interface rather than executable source.
03 / INTERACTIVE STUDY

Inspect the decision boundary

Adjust simulated model confidence, the configured threshold, and local instruction suspicion. The example assumes the model flagged an injection so the two confirmation paths are visible.

CHANGE THE INPUTS

Simulated inputs only. No page, clipboard, image, or model API is accessed. A flagged example is a review signal, not an established security finding.

ILLUSTRATIVE MODELLIVE

04 / ENGINEERING CHOICES

Why it works this way.

01

Show the evidence

Threat records retain original text, optional decoded text, suspected spans, confidence, and rationale. Review surfaces make those signals visible instead of exposing only a badge.

02

Keep page and paste policies separate

Live DOM content is highlighted in non-BLOCK modes. Clipboard surgical mode can remove spans but falls back to highlighting the original if every character would be removed.

03

Make the remote boundary explicit

Local heuristics run in the extension, while semantic and vision classification use Gemini. No separate application backend does not mean all analysis is offline.

05 / OPEN THE SOURCE

Trace it back.

Implementation details, examples, and project documentation.

Scope & limitations

  • This is a detection and review prototype, not a guarantee that a page, paste, or conversation is safe. Classifier errors and missed attacks remain possible.
  • Semantic and vision classification can send content to Gemini. Conversation monitoring also depends on supported websites’ selectors and streaming behavior.

Architecture and descriptions reflect the linked repository snapshot. The playground explains a mechanism; it does not execute the repository or report measured performance.