Files
Olaf b89d12769c
Build LinkLog Development Image / development-image (push) Successful in 28s
Privacy assesment by AI.
2026-09-30 09:08:52 +02:00

7.7 KiB

LinkLog Privacy and Fingerprinting

Scope

This document describes the privacy and fingerprinting characteristics of the LinkLog web service and Firefox extension as implemented in this repository. It is a source-code assessment, not a legal privacy policy, penetration test, or guarantee about a particular deployment.

Summary

LinkLog does not contain explicit canvas, WebGL, audio, font, hardware, timezone, analytics, or third-party advertising fingerprinting code. Its main privacy risk is different: LinkLog is designed to publish link activity, and that activity can form a highly distinctive identity profile.

A deployment operator, public visitor, upstream website, Mastodon instance, or network observer may be able to correlate a user through the data and request patterns described below.

Fingerprinting and Correlation Risks

Public activity and identity profile

The public feed exposes usernames, profile avatars, bios, exact creation timestamps, titles, original URLs, comments, tags, and whether an entry was posted to Mastodon. Public user enumeration and per-user feed URLs make it easy to collect this information for a particular account.

A sequence of saved links, topics, tags, timestamps, writing style, and referenced websites can be distinctive enough to associate a LinkLog account with activity on other services. This is a high privacy risk for users who expect saved links to be private.

Relevant implementation: backend/app/api/public.py and backend/app/services/link_service.py.

URL and query-parameter leakage

LinkLog removes a configurable list of common advertising and analytics parameters, including utm_*, gclid, fbclid, and several vendor-specific parameters. This reduces routine campaign tracking but does not make URLs anonymous.

Other query parameters, path segments, fragments that are retained by the URL cleaner, document identifiers, search terms, access codes, repository names, and user-specific URLs may still identify a person or expose sensitive information. Operators should treat saved URLs, comments, titles, and tags as potentially personal data.

The tracking-parameter list is configured through LINKLOG_TRACKING_PARAMS; changing it can also change the URL-normalization behavior of a deployment.

Firefox extension website activity

The extension uses activeTab and requests optional HTTP/HTTPS host permissions for the configured LinkLog backend. When a user saves a page, the extension reads the active page's URL and title and sends the selected data to that backend.

The extension therefore handles website activity by design. A malicious or compromised configured backend could receive the URLs that users submit, and a user can disclose sensitive page URLs by saving them. The extension stores the backend URL and username in local extension storage and keeps session credentials in Firefox session storage.

Relevant implementation: webextension/manifest.json, webextension/popup.js, and webextension/options.js.

Mastodon and external-service correlation

When enabled, the Mastodon plugin publishes the link title, comment, tags, source URL, and a UTC timestamp to the configured Mastodon instance. Posts include a recognizable User-Agent: LinkLog/1.0 in server-to-server requests. The resulting public Mastodon post can link the user's LinkLog identity, interests, and activity times to a Mastodon account.

The plugin also makes LinkLog activity observable to the configured Mastodon server, including the instance connection and publication time.

Relevant implementation: backend/app/services/plugin_manager.py.

Server-side URL fetching

The authenticated scrape endpoint fetches a user-provided URL from the LinkLog server to obtain a page title. The target website may see the LinkLog server's network address and request characteristics rather than the user's browser address. This creates server-side attribution and may reveal that a URL was submitted to LinkLog.

Deployment fingerprinting

A deployment may be distinguishable through its public version, FastAPI/OpenAPI metadata, static asset version parameters, enabled themes, feed page sizes, maximum post length, response headers, health endpoint, public hostname, and extension update metadata. These signals generally identify an installation or software version rather than a person, but they can assist cross-site correlation and targeted attack reconnaissance.

The public configuration endpoint intentionally returns feed page sizes and the maximum post-character limit. The application also exposes /health, and the README documents /docs and OpenAPI access. Production operators should decide which of these should remain public.

Authentication and account-state observation

Login throttling, response status differences, verification state, reset-mail behavior, response timing, and refresh-token behavior can reveal limited information about account state to a party able to make repeated requests. Current login errors are mostly generic, but verified and unverified account paths still differ, and failed login handling may trigger password-reset mail for known verified accounts when SMTP is configured.

Client IP handling and throttling are also deployment-sensitive. In a multi-instance deployment, the current SQLite-based limiter does not provide a shared, atomic privacy or abuse-control boundary.

Prioritize these changes for privacy-sensitive or public deployments:

  1. Make feeds and links private by default, with explicit per-link or per-profile publication controls.
  2. Make public user enumeration and profile indexing opt-in, and consider requiring authentication for user lists and private feeds.
  3. Strip or allow-list URL query parameters and warn users before saving URLs that contain credentials, access codes, search terms, or other sensitive identifiers.
  4. Offer privacy controls for avatars, bios, tags, comments, exact timestamps, and original URLs; consider timestamp coarsening for public entries.
  5. Minimize extension permissions and clearly disclose that saving a page sends its URL and title to the configured backend. Keep backend origin permissions restricted to the configured origin.
  6. Make Mastodon publication an explicit opt-in, show the complete data that will be published, and allow users to disable URL and timestamp inclusion.
  7. Protect or disable /docs, /openapi.json, detailed health/configuration endpoints, version disclosures, and unnecessary response metadata in production.
  8. Use uniform authentication responses and timing where practical, independently throttle password-reset mail, and use a shared atomic rate limiter such as Redis for multi-instance deployments.
  9. Configure strict security headers, trusted hosts, HTTPS, log retention, and access controls. Do not log authorization headers, tokens, passwords, OTP values, or secret-bearing URLs.
  10. Document retention, deletion, backup, and export behavior for links, profiles, audit records, avatars, tokens, and server logs.

User Guidance

Do not save private document links, password-reset URLs, invitation URLs, access tokens, or URLs containing sensitive query parameters to a public feed. Review titles, comments, tags, timestamps, and the Mastodon preview before publishing. Use a private deployment and disable Mastodon integration when the link history itself is sensitive.

Assessment Limits

This document reflects the repository implementation reviewed on 2026-09-16. Reverse-proxy settings, browser privacy settings, database access, server logs, deployment networking, dependencies, and third-party Mastodon behavior can materially change the effective risk. A production deployment should supplement this review with configuration review, dependency and container scanning, authenticated dynamic tests, and a retention/access-control review.