Class FCUtils

java.lang.Object
org.frontcache.core.FCUtils

public class FCUtils extends Object
  • Method Details

    • isWebComponentSubjectToCache

      public static boolean isWebComponentSubjectToCache(Map<String,Long> expireTimeMap)
    • isWebComponentCacheableForClientType

      public static boolean isWebComponentCacheableForClientType(Map<String,Long> expireTimeMap, String clientType)
    • isWebComponentExpired

      public static boolean isWebComponentExpired(Map<String,Long> expireTimeMap, String clientType)
      Check with current time if expired
      Parameters:
      clientType - {bot | guest}
      Returns:
    • dynamicCall

      public static WebResponse dynamicCall(String urlStr, Map<String, List<String>> requestHeaders, FcHttpClient client, RequestContext context) throws FrontCacheException
      GET method only for text requests for cache processor - it can use both (httpClient or filter)
      Parameters:
      urlStr -
      httpRequest -
      httpResponse -
      Returns:
      Throws:
      FrontCacheException
    • includeDynamicCallHttpClient

      public static WebResponse includeDynamicCallHttpClient(String urlStr, Map<String, List<String>> requestHeaders, FcHttpClient client, RequestContext context) throws FrontCacheException
      for includes ONLY - they allways use httpClient
      Throws:
      FrontCacheException
    • propagatedClientType

      public static String propagatedClientType(RequestContext context)
      The client type decided for the PAGE this request is a fragment of, when it can be believed. An <fc:include> re-entering a Frontcache is classified independently of the page that holds it. With a User-Agent-only bots.conf the two agree by accident - the UA survives the hop. They stop agreeing the moment a rule reads anything else: a client-ip: rule sees the sibling node's address rather than the visitor's, and a cookie: rule sees no cookies at all once the servlet request behind an async include has been recycled. Page and fragments then read different branches of the same WebResponse's expireTimeMap, and the fragments quietly stop being cached - with nothing in the log but a client-type column that disagrees between the two lines. Three conditions, all required, none of them optional:
      • the peer is trusted (front-cache.client-ip.trusted-proxies). FCUtils forwards every header it does not blacklist, so without this a client could simply send the header and choose its own cache behaviour. Unset by default, which means propagation is OFF by default and classification falls back to the rules - the safe direction;
      • the request arrived as an include (it carried a Frontcache request id). A top-level request is a client's, and a client's claim about this is worth nothing;
      • the value is a client type this node knows. Anything else is ignored rather than stored - an unknown client type would key nothing in expireTimeMap.
      Note for whoever configures the proxy: a trusted proxy that passes client headers through verbatim (nginx does, by default) should strip X-Frontcache-Client-Type - exactly as it should already be stripping X-Frontcache-Client-IP. These requests are decided BEFORE the rules run, so no bots.conf rule counts a hit for them. That is the intended accounting, and the same reason the rate limiter's scope=toplevel exists: the visitor was already classified once, when their page was requested, and counting their fragments would count them again once per fragment.
      Returns:
      the propagated client type, or null when this request must be classified on its own
    • getClientIP

      public static String getClientIP(jakarta.servlet.http.HttpServletRequest request)
    • hasMalformedQueryParams

      public static boolean hasMalformedQueryParams(String queryString)
      Detects a structurally-invalid request query string whose parameter names carry the HTML-entity artifact amp; (raw, or percent-encoded as amp%3B).

      A broken crawler that fails to decode &amp; in the origin's hrefs turns a real parameter such as &pagingPage_ci=2 into a bogus parameter named amp;pagingPage_ci, and the amp; prefix compounds (amp;amp;...) on every crawl hop. These are never legitimate parameters; worse, each unique permutation is a distinct cache key, so left unchecked they all miss the cache and flood the origin (~26% of prod traffic on fc-ap). Such requests are rejected with 400 before reaching cache or origin.

      Parameters:
      queryString - raw request query string, optionally starting with '?'; may be null/empty
      Returns:
      true if any parameter name is or starts with amp;
    • httpResponse2WebComponent

      public static WebResponse httpResponse2WebComponent(String url, FrontCacheHttpResponseWrapper originWrappedResponse, RequestContext context) throws FrontCacheException, IOException
      is used in ServletFilter mode
      Parameters:
      url -
      originWrappedResponse -
      Returns:
      Throws:
      FrontCacheException
      IOException
    • httpResponse2WebComponent

      public static WebResponse httpResponse2WebComponent(String url, org.apache.hc.core5.http.ClassicHttpResponse response, RequestContext context) throws FrontCacheException, IOException
      Throws:
      FrontCacheException
      IOException
    • transformRedirectURL

      public static String transformRedirectURL(String originLocation, RequestContext context)
      Rewrites a Location the origin sent so it names this site rather than the origin. The origin's Location is treated as a path and query with a scheme hint: its host is always discarded (the origin knows itself by an internal name - the app sees Host: origin.example.com because buildRequestHeaders drops the client's Host) and its port always comes from our own config. Three things changed here in docs/archive/guard-redirect-url-host-fix.md, all of them silent before:
      • the substituted host is the site's public name, not the Host this request arrived under - same distinction as GuardAction.publicHost, same reason;
      • the port is omitted when it is the default for the scheme. It used to be appended unconditionally, so every app redirect on a normal deployment went out as https://www.example.com:443/...;
      • an absolute origin-host URL sitting INSIDE the query string (a return= / next= parameter) is rebased too. Rewriting only the outer host is what let the origin's name reach browsers through the one part of the Location nobody was looking at.
    • revertHeaders

      public static Map<String, List<String>> revertHeaders(org.apache.hc.core5.http.Header[] headers)
      revert header from HttpClient format (call to origin) to Map
      Parameters:
      headers -
      Returns:
    • revertHeaders

      public static Map<String, List<String>> revertHeaders(jakarta.servlet.http.HttpServletResponse response)
    • convertHeaders

      public static org.apache.hc.core5.http.Header[] convertHeaders(Map<String, List<String>> headers)
    • getRequestURL

      public static String getRequestURL(jakarta.servlet.http.HttpServletRequest request)
      Parameters:
      request -
      Returns:
    • isGzipped

      public static boolean isGzipped(String contentEncoding)
      return true if the client requested gzip content
      Parameters:
      contentEncoding -
      Returns:
      true if the content-encoding containg gzip
    • buildWebComponent

      public static WebResponse buildWebComponent(String url, byte[] content, Map<String, List<String>> headers)
      Builds a WebResponse from a body and a header map, applying the same x-frontcache-component-* reading every origin response goes through. The public door onto parseWebComponent(String, byte[], Map), for the one caller that has a fragment and its headers but no HTTP response to hand: the include processor unpacking a combined response (docs/archive/combine-reduce-proposal.md section 6). Routing it here rather than letting that code assemble a WebResponse itself is what keeps maxAge, tags, refresh type and cache level meaning exactly the same for a combined member as for a single include - the invariant in section 4.2, which is silent when broken.
      Parameters:
      headers - the part's headers; a case-insensitive map is expected (header names are case-insensitive per RFC 7230) and is what the caller is given
    • maxAgeStr2Int

      public static long maxAgeStr2Int(String maxAgeStr)
    • buildRequestURI

      public static String buildRequestURI(jakarta.servlet.http.HttpServletRequest request)
    • buildRequestURI

      public static String buildRequestURI(String urlStr)
      http://localhost:8080/coin_instance_details.htm? -> /coin_instance_details.htm?
      Parameters:
      urlStr -
      Returns:
    • canonicalizeURL

      public static String canonicalizeURL(String urlStr)
      The URL a request for urlStr will actually be made with: scheme and host untouched, path and query normalised exactly as buildRequestURI(String) normalises them for the wire. This is the only correct cache key for a URL that is fetched through buildRequestURI, and the two must not be allowed to drift apart. buildRequestURI re-serializes the query and silently drops any parameter with no value - splitQueryParameters maps id= to a null value, because its guard is pair.length() > idx + 1, which for "id=" is 3 > 3. So a URL keyed raw and fetched normalised is a cache entry that can never be read back:
        probe   ?locale=zh&id=   miss - nothing is ever stored under this
        fetch   ?locale=zh         the empty parameter is gone by the time it reaches the wire
        store   ?locale=zh         what the serving node receives, and stores
        probe   ?locale=zh&id=   miss again, for the life of the entry
      
      That loop ran in production: 33 fragment URLs generating 608 origin fetches a second, each one re-fetching a fragment that was already cached, unexpired and cacheable - because the key being looked up had never existed. See docs/archive/input-request-count-fix.md section 3. Idempotent: canonicalizing an already-canonical URL returns it unchanged, so it is safe to apply on a path that may already have applied it.
    • rebaseURL

      public static String rebaseURL(String urlStr, String baseURL)
      The same path and query, addressed at a different scheme://host[:port]. This is what splits an include's two jobs apart (docs/archive/include-fetch-fix.md section 5). A cache-missing <fc:include> is keyed on the public URL and fetched from front-cache.origin-host; the two differ only in the base, so the fetch URL is derived from the public one rather than re-concatenated from the tag - re-concatenating would give a second string to keep in step with canonicalizeURL(String), and that drift is what section 2 of docs/archive/input-request-count-fix.md was. The base is expected without a trailing slash, which is what FrontCacheEngine.makeURL and URL.toString() produce for a host-only URL.
      Parameters:
      urlStr - absolute URL, or a path (which is returned appended to the base)
      baseURL - scheme://host[:port]
    • splitQueryParameters

      public static Map<String, List<String>> splitQueryParameters(String paramsStr) throws UnsupportedEncodingException
      Throws:
      UnsupportedEncodingException
    • keepAliveStrategy

      public static org.apache.hc.client5.http.ConnectionKeepAliveStrategy keepAliveStrategy()
      The keep-alive strategy shared by the engine's origin client and FrontCacheClient: honour the origin's own Keep-Alive: timeout=N, defaulting to 10s when it does not say. Extracted during the HttpClient 5 migration - the same anonymous class was pasted into FrontCacheEngine and FrontCacheClient (and into FrontCacheAgent, which cannot share it because the agent must not depend on frontcache-core). 5.x also replaced the hand-rolled BasicHeaderElementIterator with MessageSupport.iterate().
    • getHttpHost

      public static org.apache.hc.core5.http.HttpHost getHttpHost(URL host)
    • buildRequestHeaders

      public static Map<String, List<String>> buildRequestHeaders(jakarta.servlet.http.HttpServletRequest request)
    • buildRequestHeaders

      public static Map<String, List<String>> buildRequestHeaders(jakarta.servlet.http.HttpServletRequest request, boolean normalizeAcceptEncoding)
      Parameters:
      normalizeAcceptEncoding - true to replace the client's Accept-Encoding with gzip (what every cacheable path wants - see below); false to forward the client's own, which only the bypass path does, and only when encoding passthrough is enabled
    • addPublicForwardedHeaders

      public static void addPublicForwardedHeaders(Map<String, List<String>> headers, RequestContext context)
      Adds X-Forwarded-Host / X-Forwarded-Proto naming the public site to an origin call, when front-cache.origin.forward-public-host is on. The origin cannot otherwise know what the visitor typed: isIncludedHeader(String) drops the client's Host (it has to - the connection is to the origin, not to the site), so the app sees Host: origin.example.com and every absolute URL it builds names the origin. Apps work around that one entry point at a time; these two headers let the app's own framework (Spring's ForwardedHeaderFilter, Rails' trusted proxies, …) get it right everywhere at once. Off by default, because switching it on changes every absolute URL an app so configured emits. Adds nothing when the node names no public host - a made-up value here would be worse than the origin host the app already has.
    • writeResponse

      public static void writeResponse(InputStream in, OutputStream out) throws IOException
      Streams in to out. Neither is closed or flushed here - every caller already does that in a finally. The previous version had three problems, all of them on the path that streams a bypassed origin response to the client:
      • it called out.flush() after EVERY chunk, so a 5 MB download meant thousands of flushes - a write syscall each, and a chunked-transfer boundary on the wire;
      • the buffer started at 1 KB and DOUBLED every time a read filled it, with no ceiling - allocating and abandoning 2 KB, 4 KB, ... up to megabytes while streaming from a fast origin;
      • it caught IOException INSIDE the loop and printed a stack trace. A client disconnect (broken pipe) was therefore swallowed, and the loop kept pulling the whole origin body into a socket nobody was listening to, printing a trace per kilobyte.
      The IOException now propagates: the callers' finally blocks release the pooled origin connection, and a disconnected client should end the transfer rather than be written to 5000 more times.
      Throws:
      IOException
    • isAuthorizedApiKeyHeader

      public static boolean isAuthorizedApiKeyHeader(jakarta.servlet.http.HttpServletRequest request)
      Reads the credential through ApiKey.presented(HttpServletRequest) - so the stream accepts Authorization: Bearer <api-key> like everything else, and the deprecated x-frontcache-site-key for as long as that is accepted anywhere. The DECISION stays here and stays stricter than ApiKey.isAuthorized(HttpServletRequest): a node with no API key configured refuses the stream, where it would allow a management call. That difference is long-standing and deliberate - see the ApiKey javadoc - so unifying the header did not unify it.
    • applyFallbackHeaders

      public static void applyFallbackHeaders(jakarta.servlet.http.HttpServletResponse servletResponse, WebResponse webResponse)
      Apply a fallback WebResponse's status code and headers (content type plus the non-cacheable Cache-Control/Pragma markers) to the servlet response. Used by the commands that stream the fallback body directly, so that the 503 status and no-store headers set on the fallback actually reach the client / downstream caches. Must be called before the body is written.
      Parameters:
      servletResponse - target response
      webResponse - fallback response carrying status and headers