Filtering files
RWX caches tasks based on the contents of the file system (docs on caching). If a task has ever run on a particular set of files, you'll have a cache hit if you try to run that task on the same set of files.
In most cases, a task does not require every file in your repository or workspace to run successfully. For example, installing dependencies (like with npm install) usually only requires a dependency file and a lock file (e.g., package.json and package-lock.json). Because it is not affected by any other files, it's best for your task running npm install to only ever cache miss and re-execute when the files it needs change, and likewise to be a cache hit when any other file changes. You can add a filter to each of your tasks to specify exactly which files are required for that task, which will result in more frequent cache hits.
When you provide a filter for your task, only the workspace files included in the filter will be used for generating the cache key of the task. Additionally, only the workspace files in the filter will be present on disk. This prevents you from forgetting a required file from the filter, since if you do not include it in the filter, you will get a "file not found" error when the task attempts to read that file.
We recommend getting a task working and passing before adding a filter, and only then adding a filter. When initially implementing a task, it's easy to get confused about whether a failure is due to your task command and configuration being incorrect or due to the filter being incorrect. If you get the task passing without a filter, then you know the command is correct and that any failures are due to the filter being incorrectly specified.
Example
tasks:
- key: write-files
run: |
echo $(date) | tee changes-every-time.txt
echo static file contents | tee stays-the-same.txt
cache: false
- key: hash-files
use: [write-files]
run: sha256sum *.txt
- key: hash-static-files
use: [write-files]
run: sha256sum *.txt
filter:
- stays-the-same.txt
The output of hash-files will look like this:
09ebb220c8a4997acd2b9f9f99aaf200d9861b689ffcc26a3cd10c15efead761 changes-every-time.txt
672939a0624c8c668ffa85ea007e600e3c4ba503dfcf00ca1b603b7cb5c6c56b stays-the-same.txt
The output of hash-static-files will look like the following. changes-every-time.txt does not show up at all, because it was not specified in the filter:
672939a0624c8c668ffa85ea007e600e3c4ba503dfcf00ca1b603b7cb5c6c56b stays-the-same.txt
And if you run these tasks again, write-files and hash-files will always produce a cache miss, whereas hash-static-files will always produce a cache hit.
Sharing filters across tasks
A filter only applies to the task it is defined for. If a task defines a filter, any tasks that depend on that filter will not inherit the filter and will instead have every file in the workspace visible:
tasks:
- key: write-files
run: |
echo a > a.txt
echo b > b.txt
- key: read-filter
use: write-files
run: ls . # will show only a.txt
filter: [a.txt]
- key: downstream
use: read-filter
run: ls . # will show both a.txt and b.txt
If you want to share filters across tasks, you can define a YAML alias. Filters merge nested arrays, so you can even compose aliases and other filters together:
aliases:
filter-one: &filter-one [a, b]
filter-two: &filter-two [c]
tasks:
- key: write-files
run: |
echo a > a
echo b > b
echo c > c
echo d > d
- key: example-one
use: write-files
run: ls . # will show a and b
filter: *filter-one
- key: example-two
use: write-files
run: ls . # will show a and b and c
filter:
- *filter-one
- *filter-two
- key: example-three
use: write-files
run: ls . # will show a and b and d
filter:
- *filter-one
- d
In cases where you need more dynamic reusable filters, you can also use filters supplied by $RWX_VALUES:
tasks:
- key: filter-values
run: |
echo '["a", "b"]' > $RWX_VALUES/filter-one
echo '["c"]' > $RWX_VALUES/filter-two
- key: write-files
run: |
echo a > a
echo b > b
echo c > c
echo d > d
- key: example-one
use: write-files
run: ls . # will show a and b
filter: ${{ tasks.filter-values.values.filter-one }}
- key: example-two
use: write-files
run: ls . # will show a and b and c
filter:
- ${{ tasks.filter-values.values.filter-one }}
- ${{ tasks.filter-values.values.filter-two }}
- key: example-three
use: write-files
run: ls . # will show a and b and d
filter:
- ${{ tasks.filter-values.values.filter-one }}
- d
If you want to permanently remove a file from downstream tasks, you should delete the file within the task rather than relying on filters:
tasks:
- key: write-file
run: echo a > a
- key: delete-file
use: write-file
run: rm a
- key: check
use: delete-file
run: ls . # will not show any files
Filter patterns
RWX supports filter patterns for more advanced filtering use cases. Filter patterns are similar to Bash glob patterns. The following pattern features are supported:
*: match zero or more characters within a path segment. For example,foo/*.txtwill matchfoo/a.txtandfoo/b.txt, but notfoo/c/d.txt.**: match all path segments recursively. For example,foo/**/*.txtwill matchfoo/a.txt,foo/b/c.txt,foo/d/e/f.txtand so on.{}: match any of the comma-delimited patterns specified inside the braces. For example,foo/file.{js,ts}will matchfoo/file.jsandfoo/file.ts.
If you do not use filter patterns, then filter entries will only match exactly the file or directory they specify. For example, foo.txt will match foo.txt but not bar/foo.txt. If you want to match all files in the workspace named foo.txt, use **/foo.txt.
When matching a directory, all files and subdirectories are matched by default as well. The filter foo will match foo/a.txt and foo/bar/a.txt, and so on. Therefore, foo is the same as foo/* or foo/**/*.
An empty filter (filter: []) will match nothing. If you want to match everything, either exclude the filter key entirely or specify filter: ['*'].
Filter negation
You can negate any filter entry to exclude files from the filesystem. If every filter entry is negated, then all files that do not match a negative filter will be included. Otherwise, only the files that match positive filters will be included. For example, "!abc.txt" will include all files on the filesystem except abc.txt, but the filter [foo, "!foo/abc.txt"] will include only the files in the foo directory, except for foo/abc.txt.
If multiple filter entries match a file, the last filter entry to match the file wins. Thus [foo.txt, "!foo.txt"] will exclude foo.txt, but ["!foo.txt", foo.txt] will include foo.txt.
Workspace
RWX uses the contents of the entire file system for determining the cache key.
However, the filter only applies to the workspace directory, which by default is /var/mint-workspace unless you have changed it.
If any files change outside of the workspace, such as with system dependencies or configuration, it'll always result in a different cache key. Because files outside the workspace are always included in the cache key, we recommend writing any files that you may want to exclude from downstream tasks to the workspace.
Artifact filters
When a task references artifacts produced by another task, you can filter the files in those artifacts:
tasks:
- key: produce-artifact
run: |
mkdir $RWX_ARTIFACTS/my-artifact
echo "x" > $RWX_ARTIFACTS/my-artifact/x.txt
echo "y" > $RWX_ARTIFACTS/my-artifact/y.txt
- key: read-artifact
run: ls $ARTIFACT_DIR # will only show x.txt
env:
ARTIFACT_DIR: ${{ tasks.produce-artifact.artifacts.my-artifact }}
filter:
${{ tasks.produce-artifact.artifacts.my-artifact }}: [x.txt]
workspace: []
If you need to also filter files in the workspace as well as an artifact, you can specify the workspace filter in filter.workspace.