Class IncludeProcessorBase

java.lang.Object
org.frontcache.include.IncludeProcessorBase
All Implemented Interfaces:
IncludeProcessor
Direct Known Subclasses:
ConcurrentIncludeProcessor

public abstract class IncludeProcessorBase extends Object implements IncludeProcessor
Processing URL example invalid input: '<'fc:include url="/some/url/here" />
  • Field Details

    • logger

      protected org.slf4j.Logger logger
    • START_MARKER

      protected static final String START_MARKER
      See Also:
    • END_MARKER

      protected static final String END_MARKER
      See Also:
    • INCLUDE_TYPE_SYNC

      protected static final String INCLUDE_TYPE_SYNC
      See Also:
    • INCLUDE_TYPE_ASYNC

      protected static final String INCLUDE_TYPE_ASYNC
      See Also:
    • INCLUDE_FETCH_TARGET_PROPERTY

      public static final String INCLUDE_FETCH_TARGET_PROPERTY
      front-cache.include-fetch-target - where a cache-missing include is fetched from.
      See Also:
    • INCLUDE_FETCH_TARGET_ORIGIN

      public static final String INCLUDE_FETCH_TARGET_ORIGIN
      Fetch from front-cache.origin-host and cache the answer here.
      See Also:
    • INCLUDE_FETCH_TARGET_EDGE

      public static final String INCLUDE_FETCH_TARGET_EDGE
      Pre-2.9 behaviour: fetch from this node's own public URL and let the loopback do the caching.
      See Also:
    • INCLUDE_FETCH_TARGET_DEFAULT

      public static final String INCLUDE_FETCH_TARGET_DEFAULT
      What a node with nothing configured does. See includeFetchTarget.
      See Also:
  • Constructor Details

    • IncludeProcessorBase

      public IncludeProcessorBase()
  • Method Details

    • isDirectOriginFetch

      protected boolean isDirectOriginFetch(RequestContext context)
      Does a cache-missing include go straight to front-cache.origin-host? When it does, this side of the call owns the three things the loopback used to do on its way back in - the cache write, the request-log line, and the dynamic-URL check before storing. When it does not, none of them happen here, which is what makes edge a true no-op rather than a second implementation of the same behaviour. Filter mode is included, but only because isLoopbackOriginFetch(RequestContext) makes the ownership true there - read that first, it is where the interesting part is.
    • isLoopbackOriginFetch

      protected boolean isLoopbackOriginFetch(RequestContext context)
      Will this fetch come back through THIS Frontcache? Only in filter mode, and there always: the origin is the app this filter is installed in, so an HTTP call to front-cache.origin-host necessarily re-enters the filter. (In standalone mode the origin is a different process. A chained topology puts another Frontcache in the path, but that is a different node with its own keyspace, which is not this problem.)

      Why this has to be known, and what it costs to get wrong

      A re-entered request keys its own cache write on context.getCurrentRequestURL(), which is built from the Host it arrived at. So the moment the fetch is re-addressed from the public hostname to origin-host, that node - us - starts writing the fragment under a SECOND key, while this side writes it under the public one. Two entries per fragment: exactly the failure this whole change exists to avoid, arriving through a door section 8 of docs/archive/include-fetch-fix.md did not know about. It is invisible in the obvious deployment, because origin-host usually resolves to the same name the node is reached at - the two keys coincide and nothing looks wrong. It took giving the e2e filter node an origin hostname distinct from its public one to see it: every missing include wrote two request-log lines, and CommonTests.l1l2Cache counted four cached entries where three were expected.

      The answer

      The fetch is marked x-frontcache-dynamic-request, which is the existing, node-wide signal for "serve this, do not cache it" (FrontCacheEngine.processRequestInternal gates its whole cacheable branch on it, and the combine envelope has always used it for the same reason). The re-entry then runs the app and returns its response without storing anything, and this side keeps sole ownership of the single public key - the same division of labour standalone mode gets for free. What filter mode buys is smaller than standalone's: the request still lands on this node's own front door, so it is one local hop rather than a trip out to the CDN and back.
    • includeFetchBaseURL

      protected String includeFetchBaseURL(String edgeBaseURL, RequestContext context)
      The base URL includes are FETCHED from - the origin's when this node fetches includes directly, and otherwise the edge base it has always used.
      Parameters:
      edgeBaseURL - this node's own scheme://host[:port], as passed to IncludeProcessor.processIncludes(WebResponse, String, Map, FcHttpClient, RequestContext, int)
    • hasIncludes

      public boolean hasIncludes(WebResponse webResponse, int recursionLevel)
      Does this response look like it contains an <fc:include .../> worth parsing?

      Why ISO-8859-1 and not the page's real charset

      This is a byte scan wearing a String's clothes. Latin-1 maps bytes 0x00-0xFF one-to-one onto chars U+0000-U+00FF, so with compact strings new String(bytes, ISO_8859_1) performs no charset conversion at all - it copies the byte array and labels it. What that buys is String.indexOf(int), which is a vectorized JVM intrinsic, over the raw bytes. It replaced new String(content), which decoded the whole page with the platform charset - for anything non-ASCII, expanding it into a UTF-16 char array twice the size - purely to answer a boolean and throw the result away. FrontCacheEngine calls this for EVERY cacheable response that did not come from another Frontcache, and the overwhelmingly common answer is "no includes here". Measured on a ~166 KB page, per call: 64 us -> 57 us for pure ASCII, and 133 us -> 53 us once the page contains multi-byte characters. (A hand-written Boyer-Moore-Horspool scan over the raw bytes was tried and is slower - 89 us / 79 us. It allocates nothing, but no Java loop beats the intrinsic.)

      Why it is correct

      The markers are pure ASCII, and every encoding this can receive - UTF-8, Latin-1, the other single-byte sets - is ASCII-transparent: an ASCII byte never occurs inside a multi-byte sequence. So a marker is found exactly when it is really there. It is also strictly more robust than decoding: a malformed UTF-8 sequence decodes to a replacement char and SHIFTS every later index, where the Latin-1 mapping is total and index-preserving. (A UTF-16 body would not match - but neither did the old code, which decoded as UTF-8.) One nuance: MAX_INCLUDE_LENGHT is now a distance in bytes rather than in decoded chars. For a tag of the form <fc:include url="..." /> the two are identical; where they differ the byte distance is the larger, so this is marginally the more conservative check. It only gates whether the real parse runs - parseIncludes is authoritative and applies no such limit.
      Specified by:
      hasIncludes in interface IncludeProcessor
    • mergeIncludeResponseHeaders

      protected void mergeIncludeResponseHeaders(Map<String, List<String>> outHeaders, Map<String, List<String>> includeResponseHeaders)
    • getIncludeURL

      protected String getIncludeURL(String content)
      Parameters:
      content -
      Returns:
    • getIncludeType

      protected String getIncludeType(String content)
      Parameters:
      content -
      Returns:
    • getIncludeClientType

      protected String getIncludeClientType(String content)
    • getIncludeCombineGroup

      protected String getIncludeCombineGroup(String content)
      The combine attribute's value, or null when the include is not marked combinable. combine="true" groups by the include's endpoint alone; any other value adds itself to the group key, so one endpoint can serve two logically different sets on the same page without their members being merged into one batch. See docs/archive/combine-reduce-proposal.md section 8. Absent means today's behaviour, unchanged: combining is opt-in per include, because it is a contract with the origin rather than a transparent edge optimisation.
    • callInclude

      protected WebResponse callInclude(String publicURL, String fetchURL, Map<String, List<String>> requestHeaders, FcHttpClient client, RequestContext context, String includeLevel, String includeType) throws FrontCacheException
      Resolve one include: from cache when it is there, from the origin when it is not. The two URLs are the whole of docs/archive/include-fetch-fix.md. publicURL is this node's own address for the fragment and is what the entry is KEYED on - rebasing the key would orphan every existing entry and leave two keys per fragment depending on which path filled it. fetchURL is what goes on the wire. They are the same string until front-cache.include-fetch-target=origin pulls the fetch off the node's own front door.
      Parameters:
      publicURL - cache key, request log, fallback lookup
      fetchURL - what is actually requested
      Throws:
      FrontCacheException
    • getIncludeFromCache

      protected WebResponse getIncludeFromCache(String publicURL, String fetchURL, Map<String, List<String>> requestHeaders, FcHttpClient client, RequestContext context, String includeLevel, String includeType, long start)
      The cache half of callInclude(String, String, Map, FcHttpClient, RequestContext, String, String): the entry to serve, or null when this include must go to the origin. Split out for the combine path, which has to know which members missed before it can decide what goes into a batch (docs/archive/combine-reduce-proposal.md section 4.3). Everything the single-include path did on a hit - the client-type cacheability check, soft vs regular expiration, both request-log writes - still happens here and only here, so the two paths cannot drift on what a hit means.
      Parameters:
      fetchURL - used for one thing only - the soft refresh below, which is a fetch. Everything else here keys, logs and serves on publicURL
      start - when the include began, so a hit reports the same elapsed time it always did
    • fetchIncludeFromOrigin

      protected WebResponse fetchIncludeFromOrigin(String publicURL, String fetchURL, Map<String, List<String>> requestHeaders, FcHttpClient client, RequestContext context, String includeLevel, String includeType, long start) throws FrontCacheException
      The origin half of callInclude(String, String, Map, FcHttpClient, RequestContext, String, String), reached when getIncludeFromCache(String, String, Map, FcHttpClient, RequestContext, String, String, long) found nothing serveable.

      Who caches what this returns

      Historically: nobody here. The fetch went to this node's own Host, so it arrived back at the front door as an ordinary request and CacheProcessorBase.processRequest stored it on the way through, and FrontCacheEngine wrote its request-log line. Fetch the origin directly and both of those stop happening - so both are done here instead, and only then (isDirectOriginFetch(RequestContext)). Miss the cache write and every include becomes an uncached origin call, forever, with nothing in the logs to say so: failure mode 1 of docs/archive/include-fetch-fix.md section 8.
      Parameters:
      publicURL - cache key and log line - never the URL this fetched with
      fetchURL - what goes on the wire
      start - when the include began - including the cache lookup that just missed, so the reported time is the whole include and not just its origin leg
      Throws:
      FrontCacheException
    • isIncludeResponseCacheable

      protected boolean isIncludeResponseCacheable(WebResponse webResponse, RequestContext context)
      The four conditions an include response must meet before it is stored, shared by the single include path and the combine path so the two cannot drift on what "cacheable" means. They are the same decision. isCacheable() on its own is NOT the guard, and that is the trap: it answers true for a response whose expireTimeMap says NO_CACHE for every client type, because it only checks that the map is non-empty. A response with no maxage would therefore be stored as an entry that is expired the instant it is written - churning the index and never serving anything.
    • isDynamicURL

      protected static boolean isDynamicURL(String url)
      Would dynamic-urls.conf have stopped the loopback from caching this? Matched against the URI path, which is what FrontCacheEngine.ignoreCache matches - an include cached on this side must be cacheable exactly when the loopback would have cached it.
    • logCombinedInclude

      protected void logCombinedInclude(String urlStr, RequestContext context, String includeLevel, long elapsed, long lengthBytes)
      Request-log and trace-header lines for one member of a combined batch. A single include that misses the cache is logged twice: once here as a trace header, and once by FrontCacheEngine when the loopback request for it arrives. A combined member has no loopback request of its own - that is the saving - so without this it would vanish from the request log entirely, and the log would show one origin request where fifteen fragments were served. Both lines are therefore written here.
      Parameters:
      elapsed - the BATCH's elapsed time; see FCHeaders.COMPONENT_COMBINED_INCLUDE
    • init

      public void init(Properties properties)
      Reads front-cache.include-fetch-target. A subclass that overrides this must call super.init(properties) to honour the property; one that does not keeps edge, which is what it does today.
      Specified by:
      init in interface IncludeProcessor
    • destroy

      public void destroy()
      Specified by:
      destroy in interface IncludeProcessor