Skip to content

Repository files navigation

ZIPsFS: Access ZIP files transparently as if they were regular folders

Overlay file system which expands ZIP files

Example usage

ZIPsFS [ZIPsFS-options] path-of-branch1/  branch2/  branch3/  branch4/  :  [fuse-options] mount-point

Summary

ZIPsFS functions as a union or overlay file system, merging multiple file structures into a unified directory. This directory presents the underlying files and sub-directories from the specified sources (branches) as a single, cohesive structure. Any newly created or modified files are stored in the first file location, while all other sources remain read-only, ensuring that their files are never altered. ZIPsFS treats ZIP files as expandable folders, typically naming them by appending .Contents/ to the original ZIP file name. However, this behavior can be customized using filename-based rules. Extensive configuration options allow adjustments. Changes can be applied without disrupting the file system.

With a trailing slash, the folder name is not part of the virtual path, in accordance with the trailing slash semantics of many UNIX tools. More flexibility provides the property @path-prefix=....

ZIPsFS includes specialized features like automatic file conversions and performance optimizations tailored for efficiently storing and accessing large-scale mass spectrometry data. It manages non-sequential file reading from varying positions (so-called file seek) efficiently to improve reading of remote and zipped files. It has a programming interface to create synthetic file content programmatically.

ZIPsFS is best run in a tmux session.

Mini tutorial

Please Install ZIPsFS.

Create some example files

Without trailing slash, the folder name will be retained in the virtual path. This is the case for branch3. Virtual file paths in that branch will start with mount point/branch3/.

b1=~/test/ZIPsFS/writable/
b2=~/test/ZIPsFS/branch1/
b3=~/test/ZIPsFS/branch2/
b4=~/test/ZIPsFS/branch3

mkdir -p $b1 $b2 $b3 $b4 ~/test/ZIPsFS/mnt

for c in a b c d e f; do echo hello world $c >$b2/$c.txt; done
for ((i=0;i<10;i++)); do echo hello world $i >$b3/$i.txt; done

zip --fifo $b2/zipfile1.zip <(date)  <(echo $RANDOM)
zip --fifo $b3/zipfile2.zip <(hostname)  <(ls /)
zip --fifo $b3/20250131_this_is_a_mass_spectrometry_folder.d.Zip   <(seq 10)

Start ZIPsFS

In production, it is recommended to start ZIPsFS in tmux. For testing, just use your regular command line.

ZIPsFS   $b1 $b2 $b3 $b4  :  ~/test/ZIPsFS/mnt

Browse the virtual file tree

Open a file browser or another terminal and browse the files in

~/test/ZIPsFS/mnt/

Create a file in the virtual tree

The first file tree stores files. All others are read-only.

echo "Hello" > ~/test/ZIPsFS/mnt/my_file.txt

cat ~/test/ZIPsFS/mnt/my_file.txt

To get the real storage place of the file, append @SOURCE.TXT

cat ~/test/ZIPsFS/mnt/my_file.txt@SOURCE.TXT

More information is available with suffix @PROPRERTIES.TXT

cat ~/test/ZIPsFS/mnt/my_file.txt@PROPRERTIES.TXT

Access web resources as regular files (Not fully tested yet)

Make sure the UNIX tool curl is installed.

curl --version

Note that "//:" and all slashes in the URL are replaced by commas.

less  ~/test/ZIPsFS/mnt/zipsfs/n/ftp,,,ftp.uniprot.org,pub,databases,uniprot,LICENSE

gunzip -c ~/test/ZIPsFS/mnt/zipsfs/n/ftp,,,ftp.ebi.ac.uk,pub,databases,uniprot,current_release,knowledgebase,complete,docs,keywlist.xml.gz

Even though the file is only available in compressed form on the server, you can directly access the decompressed file. Omit the .gz extension.

The decompressed file size is an estimate. It becomes exactly known after reading the file.

head ~/test/ZIPsFS/mnt/zipsfs/n/http,,,ftp.uniprot.org,pub,databases,uniprot,current_release,knowledgebase,reference_proteomes,Eukaryota,UP000005640,UP000005640_9606.fasta
head ~/test/ZIPsFS/mnt/zipsfs/n/ftp,,,ftp.uniprot.org,pub,databases,uniprot,current_release,knowledgebase,reference_proteomes,Eukaryota,UP000005640,UP000005640_9606.fasta


ls -l ~/test/ZIPsFS/mnt/zipsfs/n/ftp,,,ftp.ebi.ac.uk,pub,databases,uniprot,current_release,knowledgebase,complete,docs,keywlist.xml

head ~/test/ZIPsFS/mnt/zipsfs/n/ftp,,,ftp.ebi.ac.uk,pub,databases,uniprot,current_release,knowledgebase,complete,docs,keywlist.xml

From now on, the file size is known.

The gz compression is built-in using the zlib C library. The bz2, xz and lrz decompression work only if activated in the config (default) and if the respective backends are installed on the computer. The following are test examples of files remotely stored as bz2 and xz, respectively.

head ~/test/ZIPsFS/mnt/zipsfs/n/http,,,ftp.gnu.org,gnu,binutils,binutils-2.23.1.tar   # Remote file is .tar.bz2
head ~/test/ZIPsFS/mnt/zipsfs/n/http,,,ftp.gnu.org,gnu,gawk,gawk-4.0.2.tar            # Remote file is .tar.xz

Above method allows to access any remote file. The disadvantage over the method below is, that repositories are not browsable.

Browsing FTP sites using nested FUSE file systems

The script file ZIPsFS_prepare_branch_for_ftp.sh creates folders like ~/.ZIPsFS/db/pride and mounts the respective FTP sites. This requires curl. Please also install rclone (or curlftpfs).

Its standard output serves as CLI parameters for the command line of ZIPsFS.

ZIPsFS   $b1 $b2 $b3 $b4  $(./ZIPsFS_prepare_branch_for_ftp.sh)   :  ~/test/ZIPsFS/mnt

Check the repositories:

ls ~/test/ZIPsFS/mnt/db

GZ compressed files are transparently de-compressed. Files with the ending .gz also appear in the file listing without gz suffix.

But wait, there is a major problem. Before decompression, ZIPsFS can only communicate an estimated file size of the decompressed file. Consequently, programs that depend on exact file sizes will not work. The command /usr/bin/tail requires the length of the file to find the last lines. Please try

tail  ~/test/ZIPsFS/mnt/mnt/db/pride/2005/08/PRD000004/PRIDE_Exp_Complete_Ac_369.xml

The first time this command is run, there is no output. This is because tail sees the wrong file size. Subsequently, tail works nicely.

You can remove the preloaded file and again tail will fail once.

rm  ~/test/ZIPsFS/mnt/mnt/db/pride/2005/08/PRD000004/PRIDE_Exp_Complete_Ac_369.xml

DESCRIPTION

ZIPsFS is a Union / overlay file system

ZIPsFS functions as a union (overlay) file system. When files are created or modified, they are stored in the first file tree. If a file exists in multiple source locations, the version from the leftmost source (the first one listed) takes precedence. With an empty string as the first source, the ZIPsFS file system read-only and file creation and modification is disabled.

The physical file path, i.e., the actual storage location of a file, can be retrieved from a file formed by appending @SOURCE.TXT to the filename. For example, to determine the real location of: ~/test/ZIPsFS/mnt/1.txt Run the following command:

cat ~/test/ZIPsFS/mnt/1.txt@SOURCE.TXT

If a remote file branch system stops responding, the current file access is blocked (Unless the options WITH_ASYNC_READDIR ... are activated). After some time, unresponsive file trees will be skipped to prevent that file accesses get blocked.

ZIPsFS expands ZIP file entries

By default, ZIP files are displayed as folders with the suffix .Content. This behavior can be customized. The default configuration includes a few exceptions tailored to specific use cases in Mass Spectrometry:

  • For ZIP files whose names end with .d.Zip, the virtual folder will end with .d.

  • Flat file list: For Sciex instruments, each mass spectrometry record is stored in a set of files which are not organized in sub-folders. For example the record 20231122_MA_HEK_QC_hiFlow_2ug_SWATH_rep01 stored in 20231122_MA_HEK_QC_hiFlow_2ug_SWATH_rep01.wiff2.Zip may consist of the following files

    • 20231122_MA_HEK_QC_hiFlow_2ug_SWATH_rep01.timeseries.data
    • 20231122_MA_HEK_QC_hiFlow_2ug_SWATH_rep01.wiff
    • 20231122_MA_HEK_QC_hiFlow_2ug_SWATH_rep01.wiff2
    • 20231122_MA_HEK_QC_hiFlow_2ug_SWATH_rep01.wiff.scan

    These are contained with those of other entries in the file listing. To get the full list of files, all ZIP files need to be inspected. This is time consuming and performed in the background. Consequently, the file listing will be incomplete when requested for the first time. Only after some seconds, the file listing will be presented properly.

A major advantage of using ZIP archives for storing files is that Zip files contain a record of the CRC32 checksum of each zip entry. This can be used for integrity checks to detect data decay or problems with data transfer. This checksum can be obtained from ZIPsFS by appending @ARCHIVECRC32.TXT to a file path.

Command line Options

-h

Prints brief usage information.

*-s path-of-symbolic-link This is discussed in section Configuration.

-c [NEVER,SEEK,RULE,COMPRESSED,ALWAYS]

Policy when ZIP entries and file content is cached in RAM.

NEVER ZIP entries are never cached, even not in case of backward seek.
SEEK ZIP entries are cached when the file position jumps backward.
RULE ZIP entries are cached according to customizable rules or on backward seed. (DEFAULT)
COMPRESSED All compressed ZIP entries are cached.
ALWAYS All ZIP entries are cached.

-l Maximum memory for caching ZIP-entries in the RAM

Specifies a limit for the cache. For example -l 8G would limit the size of the cache to 8 Gigabyte. File content larger than this will not be cached. When memory usage is high, cached file access waits until it drops below this value.

-b Execution in background (Not recommended). We recommend running ZIPsFS in foreground in tmux.

These rules can be overridden by using the sub-directories /zipsfs/-/m/ and /zipsfs/-/-m/. Please see enclosed README files.

FUSE command line options

Options for the FUSE system come after the colon in the command line.

-o comma separated Options

-o allow_other Other users are granted access.

The last argument is the mount point which is an empty folder.

Special directory /zipsfs/

The folder /zipsfs/ provides alternative views and file prefetch. It contains log files and scripts. The file tree can be accessed in a modified way via

/zipsfs/<view-option>/<file preload-option>[/<preload-selector>]

These three folders contain single letter directives which take precedence over root-folder settings at the command line.

Folder /zipsfs/<view-option>/

The 1st level sub-directories specifies view options. Directives are single letters. They can be combined. A preceding dash '-' negates.

  • /zipsfs/- Default view.
  • /zipsfs/z Rapid navigation and file name searching without time consuming ZIP file expansion.
  • /zipsfs/n Internet files. Take URL and replace colon and slashes by comma. See above tutorial.
  • /zipsfs/d For bz2, xz, gz, .Z and lrz compressed files also generate the decompressed file.
  • /zipsfs/1 Files of the first branch only. This includes all files ever written or generated or modified.
  • /zipsfs/~1 Files except from first branch.
  • /zipsfs/c Show converted files along the original files. Example: - Raw mass spectrometry to mgf or mzML files or tsv - Parquet to tsv
  • /zipsfs/log Logging, to identify very busy software which should rather be used with preloading of remote or compressed files.
    • Excessive requests of file attributes
    • Multiple open/close
    • Backward seek
    • Upper/lower case conversion of file names

Details are found in the contained readme files.

Folder /zipsfs/<view-option>/<prefetch option>

The following path component may instruct preloading of files. Files are preloaded just before the first byte is read.

  • /-/ Apply default preload rules defined in ZIPsFS_configuration.c and ZIPsFS_configuration.h and specified by root properties.
  • /m/ Prefetch to RAM
  • /-m/ Do not prefetch to RAM
  • /l/ Prefetch to the local file system
  • /lz/ Prefetch to the local file system. Load entire ZIP files not their entries. Do not mix up with /l/z. In the later, z is the selector "If-in-ZIP".
  • /-l/ Do not prefetch to the local file system

The /m/ and /l/ folder have a subfolder with selectors: (a) all, (r) remote and (z) ZIP-entry.

Properties of file branches

In the command line, root file paths can be followed by expressions like @immutable=1 to set specific properties. Alternatively, properties can be given in a file <root-path>.ZIPsFS.properties. This is demonstrated in ZIPsFS_prepare_branch_for_ftp.sh.

For a list of properties run

ZIPsFS -h
Preload files

Non-linear file loading in a random-access manner is very common for Windows style software were pipes and process substitution are rarely used. This is also the case for the widely used mass-spectrometry programs Diann and FragPipe (Fragger). Vendor specific mass-spectrometry files are loaded from varying file positions. Apparently, this mode of file access is inefficient for remote or ZIP-compressed data.

File content pre-load

Data preloading can help to provide file data out-of-order. Remote or compressed files can be preloaded in two different ways: (I) into RAM or (II) to the local disk.

Preloading into RAM

File names to be cached in RAM are specified in the configurable method config_advise_cache_in_ram(). The default setting includes Brukertimstof mass-spectrometry files. Preloading to the RAM is appropriate for these files because each file is loaded only once per analysis. The -l option sets an upper limit on memory usage for the ZIP RAM cache. When available memory runs low, ZIPsFS can either pause, proceed without caching file data or just ignore the memory restriction depending on the configuration.

Preloading to local disk

File reading of remote or compressed files can be improved by caching the file content on the local disk. This is for example necessary for Thermo mass-spectrometry raw files analyzed in FragPipe. Above method of preloading into RAM is here inappropriate because each raw file is opened and closed multiple times during computation such that the large files would be transfered over the net several times.

All remote (r) or compressed (c) or zippded (z) files accessed through the following folders will be first copied to local disk:

<mount-point>/zipsfs/-/l/r/
<mount-point>/zipsfs/-/l/z/
<mount-point>/zipsfs/-/l/rz/

The Readme in these folders provide further information. Alternatively, the entire root-path can be marked for preloading with the property @preload.

Preloading complete ZIP files to the local disk - a combination of both

With the branch

<mount-point>/zipsfs/-/l/r/

each ZIP entry would be preloaded separately.

However, with

<mount-point>/zipsfs/-/lz/r/

the ZIP file will be preloaded. Depending on the settings, the ZIP entries might be preloaded to RAM.

Transiently caching file attributes and directory listings

When loading Bruker mass spectrometry files with the Bruker DLLs, The file system is asked douzens of times per second to provide one and the same file listing and the same file attributes. While this is no problem for local files, remote file systems get overloaded. ZIPsFS implements a transient cache, that lives as long as the specific ZIP file is used.

This cache is distinct from a persistent cache of directory listing that may last for ever for root branches with the property @immutable.

Decompression

Decompression is enabled with properties like @preload=1 of for specific compression suffixes like @preload=gz,xz or for the file tree

<mount-point>/zipsfs/d/
Project status

Author: Christoph Gille

Current status: Testing and Bug fixing. Occasionally improving concepts. Already running very busy for several weeks without interruption.

If ZIPsFS crashes, please send the stack-trace together with the source code you were using.

End ZIPsFS

ZIPsFS can be killed like any UNIX command by typing Ctrl-C. Sometimes this does not work. Assuming mnt is your mount-point, the following might work:

fusermount -u mnt
sudo umount mnt

If ZIPsFS still hangs, the following will kill all FUSE file systems of the current user. Those of other users won't be affected:

echo 1 | sudo tee /sys/fs/fuse/connections/*/abort

Changelog

  • 202601
    • Changing ZIPsFS to lower-case in /zipsfs. This was necessary to export ZIPsFS as a Samba Share to Windows
  • 202610
    • Re-organized the /zipsfs file branch.
    • Solved problem of file truncation due to underestimated file size
Configuration ZIPsFS can be customized:
  • optional features can be (de)-activated with preprocessor macros "WITH_SOME_FEATURE" which take the values 0 or 1.
  • Rules can be given
    • which files are cached in the main memory
    • which ZIP entries are inlined
  • Timeout values for accessing (remote) files
  • Automatic file conversions

Configuration files of ZIPsFS, are files written in the programming language C. They have the prefix ZIPsFS_configuration and the suffix .h of .c.

ZIPsFS is customized for our needs - accessing archived high throughput data such that it can be directly used for mass spectrometry software. These settings can be used as a sample to customize it for other needs.

Changes require recompilation and take effect after restart of ZIPsFS.

With the -s option, the updated ZIPsFS can seamlessly replace running instances without disrupting the virtual file system.

To illustrate how this works, let MNT represent the apparent mount point of the FUSE file system. Suppose we are in the parent directory of MNT, enabling the use of relative paths. Users access files through this apparent mount point, but in reality, MNT is a symbolic link to the actual mount point. The real mount point is not directly accessed by users, as it changes each time a new instance of ZIPsFS is launched.

For example, assume the obsolete ZIPsFS instance is mounted at ./.mountpoints/MNT/1. When a new instance replaces it, it may use any empty directory as mount point. ZIPsFS must be started with the following command line option:

-s MNT

Once the new instance is running, the symbolic link is updated to point to the new mount location. From the user's perspective, nothing changes - the apparent mount point remains MNT. To ensure uninterrupted access, the obsolete instance should remain active for a short period to allow ongoing file operations to complete.

If MNT is within an exported SAMBA or NFS path the real mount points should be in the exported file tree as well. Include into /etc/samba/smb.conf:

follow symlinks = yes
Generated (synthetic) files: Automatic file conversions, Accessing web resources as regular files Computations often require files from public repositories. Files from the internet (http, ftp, https) can be accessed as files using the URL as file name. ZIPsFS takes care of downloading and updating. They are immutable and cannot be modified unintentionally. In DOS, a trailing colon is a signature for device names. Therefore, the colon and all slashes in the URL need to be replaced by comma. Comma has been chosen as a replacement because it normally does not occur in URLs. Furthermore, it does not require quoting in UNIX shells.

Example with mnt/ denoting the mountpoint of the ZIPsFS file system:

sudo apt-get install curl
ls -l  mnt/ZIPsFS/n/https,,,ftp.uniprot.org,pub,databases,uniprot,README
more   mnt/ZIPsFS/n/https,,,ftp.uniprot.org,pub,databases,uniprot,README
head   mnt/ZIPsFS/n/https,,,ftp.uniprot.org,pub,databases,uniprot,README@SOURCE.TXT

To see the real local file path append @SOURCE.TXT to the file path.

The http-header is updated according to a time-out rule in ZIPsFS_configuration.c. Whether the file itself needs updating is decided upon the Last-Modified attribute in the http or ftp header.

Additionally, the file is accessible with a file-name containing the time reported in the header. This feature can be conditionally deactivated.

Generation of files using programming language C

By modifying the file ZIPsFS_configuration_c.c, users can easily implement file generation using the programming language C.

This path gives the output of the included minimal example.

<mount point>/example_generated_file/example_generated_file.txt

File conversion rules

ZIPsFS can generate virtual files by file conversion. This feature is enabled by setting the preprocessor macro WITH_FILECONVERSION to 1 in ZIPsFS_configuration.h. Generated files are stored in the first file branch, allowing them to be served instantly upon repeated requests. The default rules, defined in ZIPsFS_configuration_fileconversion.c, include:

  • Image files (JPG, JPEG, PNG, GIF): Smaller versions at 25% and 50% scaling.
  • Image files (OCR): Extracted text using Optical Character Recognition (OCR).
  • PDF files: Extracted ASCII text.
  • ZIP files: Consistency check reports, including checksums.
  • Mass spectrometry files: mgf (Mascot) and mzML formats.
  • wiff files: Extract ASCII text.
  • Apache Parquet files: TSV and TSV.BZ2 formats.

For testing, you can copy an image file (png,jpg) and find the downscaled version in the branch

/ZIPsFS/c/

Note that some of the conversions may require Docker support and membership of the docker group.

Handling Unknown File Sizes in Virtual File Systems

The system cannot determine the size of files whose content has not yet been generated. In kernel-managed virtual file systems such as /proc and /sys, virtual files typically report a size of zero via stat(). Despite this, they are not empty and contain dynamically generated content when read.

However, this behavior does not translate well to FUSE-based file systems.

For FUSE, returning a file size of zero to represent an unknown or dynamic size is not recommended. Many programs interpret a size of 0 as an empty file and will not attempt to read from it at all. In ZIPsFS, a placeholder or estimated size is returned if the file content has not been generated. The estimate should be large enough to allow reading the full content. If the size is underestimated, data may be read incompletely, leading to truncated output or application errors. This workaround allows programs to read the file as if it had content, even though the size isn’t known in advance. However, it may still break software that relies on accurate size reporting for buffering or memory allocation.

Example Diann: The mass-spectrometry program Diann does not rely on correct file sizes when reading fasta files. Since Diann does not support fasta.gz files directly, it can benefit from ZIPsFS capability to decompress on download.

Example Fragpipe: Fragpipe is a software to process mass-spectrometry files. Processing Thermo-Fisher mass-spectrometry files with the suffix raw, those are converted by Fragpipe into the free file format mzML. Since ZIPsFS can also convert raw files to mzML, we tried to give the virtual mzML files as input. Initially, their reported file size is 99,999,999,999 Bytes. This large number was chosen to make sure that the estimated file size is larger than the real yet unknown size. Initially Fragpipe attempts to read some bytes from the end of the file. To determine the reading position, it uses the overestimated file size. In this specific case it tried to read at file position 99,999,997,952. ZIPsFS will perform the conversion when serving the first read request. Since the converted mzML file is much smaller than the read position, there will be no data and Fragpipe will fail. When however, at least one byte of the mzML files is read to initiate the conversion process before Fragpipe is started, computation will succeeds.

Logging ZIPsFS typically runs as a foreground process. To keep it active and monitor its output, it is recommended to use a persistent terminal multiplexer such as tmux. This enables continuous observation of all messages and facilitates long-running sessions. Additional log files are stored in:
~/.ZIPsFS

For each mount point there are files specifying more logs.

log_flags.conf

See readme for details:

log_flags.conf.readme

ZIPsFS dynamically generates a HTML status file in /ZIPsFS/ with real-time information.

Fault management. Timouts for remote upstream file systems. Duplicated remote trees Accessing remote files inherently carries a higher risk of failure. Requests may either:
  • Fail immediately with an error code, or

  • Block indefinitely, causing potential hangs.

In many FUSE file systems, a blocking access can render the entire virtual file system unresponsive. ZIPsFS addresses this with built-in fault management for remote branches.

Remote roots in ZIPsFS are specified using a double-slash prefix, similar to UNC paths (//server/share/...). Each remote branch is isolated in terms of fault handling and threading and has its own thread pool, ensuring faults in one do not affect others. To avoid blocking the main file system thread, remote file operations are executed asynchronously in dedicated worker threads.

Timeouts

Timeouts are activiated with the respective macro definitions WITH_ASYNC_XXXXXX. The work for root-paths starting with tripple slash. The fuse thread delegates the file operation to another thread and waits for its completion. After the configurable timeout it gives up.

Timeout is not fully tested yet. It should be considered only for remote paths when blocks are observed.

The worker thread may block permanently. In this case it may be killed automatically and restarted. However killing this thread sometimes does not work.

If the stalled thread cannot be terminated, ZIPsFS will not create a new thread. To check whether all threads are responding, activate logging. For details see

~/.ZIPsFS/.../log_flags.conf.readme

This problem is best resolved by replacing the current by a new ZIPsFS session using the -s command line option.

Duplicated file trees

For redundantly stored files (i.e., available on multiple branches), another branch may take over transparently if one fails or becomes unresponsive.

Debug Options

The ZIPsFS option -T

Checks whether ZIPsFS can generate and print a backtrace in case of errors or crashes. This feature elies on external tools to translate memory addresses into source code locations: On Linux and FreeBSD, it uses addr2line, typically located in /usr/bin/. On macOS, it uses the atos tool instead. Ensure these tools are installed and accessible in your system's PATH for backtraces to work correctly.

See ZIPsFS.compile.sh for activation of sanitizers.

File integrity checks For ZIP entries loaded entirely into RAM: ZIPsFS performs CRC checksum validation. Any detected inconsistencies are logged, helping to detect corruption or transmission errors.
Use case - Archive of high throughput mass spectrometry files

We use closed-source proprietary Windows software to read large experimental data from various types of mass spectrometry machines. The data is immediatly copied into an intermediate storage on the processing PC and eventually archived in a read-only WORM file system.

To reduce the number of individual files and disk usage and to allow for data integrity checks, all files from a single mass spectrometry measurement are bundled into one ZIP archive. With fewer individual files, searching through the entire directory hierarchy takes less than 1 hour.

We initially hoped that files inside ZIP archives would be accessed using

  • Pipes
  • Named pipes
  • Process substitution
  • FUSE file systems with which transparently expand multiple ZIP files
  • Unzipping and storing extracted files on disk

Unfortunately, these techniques did not work for our use case. Mounting individual ZIP files was initially the only solution. But when sample size of large experiments got large, even this was not feasable.

ZIPsFS was developed to solve the following problems:

  • Growing Number of ZIP Files: Recently, the size of our experiments - and therefore the number of ZIP files - has increased enormously. Mounting thousands of individual ZIP files results in a very long /etc/mtab file and puts a significant strain on the operating system.

  • Write Access Requirements: Some proprietary software requires write access to both files and their parent directories.

  • Inefficiency in Random File Access: Some mass spectrometry files are read from varying positions. Random access is particularly inefficient for compressed ZIP entries, in particular with backward seeks. Buffering of file content is required.

  • Multiple Storage Locations: Experimental records are initially stored in an intermediate storage location and, after verification, are moved to the final archive. Consequently, we need a union file system.

  • Resilience of storage systems: Sometimes access to the archive gets blocked. Otherwise there are several alternative entry points which will continue to work. This adds requirement for fault management.

  • Redundant File System Requests: Some proprietary software generates millions of redundant requests to the file system, which is problematic for both remote files and mounted ZIP files. File attributes need to be cached.

File tree with zip files on hard disk:
 ├── src
 │   ├── InstallablePrograms
 │   │   └── some_software.zip
 │   │   └── my_manuscript.zip
 └── read-write
    ├── my_manuscript.zip.Content
            ├── my-modified-text.txt
       
Virtual file tree presented by ZIPsFS:
 ├── InstallablePrograms
 │   ├── some_software.zip
 │   └── some_software.zip.Content
 │       ├── help.html
 │       ├── program.dll
 │       └── program.exe
 │   ├── my_manuscript.zip
 │   └── my_manuscript.zip.Content
 │       ├── my_text.tex
 │       ├── my_lit.bib
 │       ├── fig1.png
 │       └── fig2.png
       
The file tree can be adapted to specific needs by editing ZIPsFS_configuration.c. Our mass-spectrometry files are processed with special software. It expects a file tree in its original form i.e. as files would not have been zipped. Furthermore, write permission is required for files and containing folders while files are permanently stored and cannot be modified any more. The folder names need to be ".d" instead of ".d.Zip.Content". For Sciex (zenotof) machines, all files must be in one folder without intermediate folders.
File tree with zip files on our NAS server:
 ├── brukertimstof
 │   └── 202302
 │       ├── 20230209_hsapiens_Sample_001.d.Zip
 │       ├── 20230209_hsapiens_Sample_002.d.Zip
 │       └── 20230209_hsapiens_Sample_003.d.Zip

...

│ └── 20230209_hsapiens_Sample_099.d.Zip └── zenotof └── 202304 ├── 20230402_hsapiens_Sample_001.wiff2.Zip ├── 20230402_hsapiens_Sample_002.wiff2.Zip └── 270230402_hsapiens_Sample_003.wiff2.Zip ... └── 270230402_hsapiens_Sample_099.wiff2.Zip

Virtual file tree presented by ZIPsFS:
 ├── brukertimstof
 │   └── 202302
 │       ├── 20230209_hsapiens_Sample_001.d
 │       │   ├── analysis.tdf
 │       │   └── analysis.tdf_bin
 │       ├── 20230209_hsapiens_Sample_002.d
 │       │   ├── analysis.tdf
 │       │   └── analysis.tdf_bin
 │       └── 20230209_hsapiens_Sample_003.d
 │           ├── analysis.tdf
 │           └── analysis.tdf_bin

...

│ └── 20230209_hsapiens_Sample_099.d │ ├── analysis.tdf │ └── analysis.tdf_bin └── zenotof └── 202304 ├── 20230402_hsapiens_Sample_001.timeseries.data ├── 20230402_hsapiens_Sample_001.wiff ├── 20230402_hsapiens_Sample_001.wiff2 ├── 20230402_hsapiens_Sample_001.wiff.scan ├── 20230402_hsapiens_Sample_002.timeseries.data ├── 20230402_hsapiens_Sample_002.wiff ├── 20230402_hsapiens_Sample_002.wiff2 ├── 20230402_hsapiens_Sample_002.wiff.scan ├── 20230402_hsapiens_Sample_003.timeseries.data ├── 20230402_hsapiens_Sample_003.wiff ├── 20230402_hsapiens_Sample_003.wiff2 └── 20230402_hsapiens_Sample_003.wiff.scan

...

       ├── 20230402_hsapiens_Sample_099.timeseries.data
       ├── 20230402_hsapiens_Sample_099.wiff
       ├── 20230402_hsapiens_Sample_099.wiff2
       └── 20230402_hsapiens_Sample_099.wiff.scan
Implementation

Dependencies

  • libfuse
  • libzip
  • Linux or UNIX
  • Gnu-C or clang

Language

For optimal performance, ZIPsFS is written in stanmdard C.

Roots

When a client program requests a file from the virtual FUSE file system, the virtual file path needs to be translated into an existing real file or ZIP entry.

All given roots are iterated until the file or ZIP entry is found. The first root is the only that allows file modification and file creation. All others are read only.

Caches

File attribute and file content caches in ZIPsFS

There are several caches for file attributes and file contents to improve performance. They can be deactivate using conditional compilation. The respective macro definitions and detailed description are found in ZIPsFS_configuration.h.

The caches are necessary because

  • Some mass spectrometry software reads files not sequentially but from varying file positions.
  • Mass spectrometry software excerts thousands and millions of redundand file requests
  • One type of mass spectrometry instruments requires the ZIP entries to be displayed flat in the folder. To list this directory sufficiently fast requires a file entry cache.

The file cache of the OS

After closing a file or disposing the memcache of a ZIP entry, a call to posix_fadvise(fd,0,0,POSIX_FADV_DONTNEED); may remove the file from the file cache of the OS when it is unlikely that the same file will be used in near future. This can be customized in ZIPsFS_configuration.c.

The cache in libfuse

See field struct fuse_file_info->keep_cache in function xmp_open().

Infinite loops running in own threads

  • infloop_memcache(): Loading entire ZIP entries into RAM asynchronously. One thread per root.
  • infloop_async(): Calling stat asynchronously. Reading file attributes and directory listings and ZIP file entry listings asynchronously. One thread per root.
  • infloop_misc(): Periodically check whether roots are responsive. Unblock blocked threads by calling pthread_cancel(). They re-start automatically with a pthread_cleanup_push() hook. One global thread.

When a source file system becomes unavailable

Typically, overlay or union file systems are affected when one of the source file systems is not responding. Remote sources are at risk of failure when for example network problems occur.

Those paths of roots can be indicated by a leading double slash // (according to a convention from MS Windows). Those roots are used only when they have successfully returned from periodical calls to fsstat().

The worst case is that a call to the file API does not return. Optionally, ZIPsFS allows such calls to be run in a separate thread. This is implemented with a queue. Periodically, an infinite loop in a separate thread processes the requests in the queue.

With three slashes, asynchronous calls to stat(), readdir() and open(). This requires that the respective WITH_XXX macro is set to 1.

Only read(int fd, void *buf, size_t count) and the equivalent for ZIP entries are called directly suche that some risk of blocking remains.

Limitations and Bugs There are some Limitations:

Hard Links

Hard links are not supported, though symlinks are fully functional.

Deleting Files

Files can only be deleted if their physical location resides in the first source. Files located in other branches are accessed in a read-only mode, and deletion of these files would require a mechanism to remove them from the system, which is currently not implemented.

If you require this functionality, please submit a feature request.

Reading and Writing

Simultaneous reading and writing of a file using the same file descriptor will only function correctly for files stored in the writable source.

File size

When file sizes are not known, an over-estimate will be reported. If the estimate is too low, the read file content will be truncated.

See also related sites - https://github.com/openscopeproject/ZipROFS - https://github.com/google/fuse-archive - https://bitbucket.org/agalanin/fuse-zip/src - https://github.com/google/mount-zip - https://github.com/cybernoid/archivemount - https://github.com/mxmlnkn/ratarmount - https://github.com/bazil/zipfs

About

Accessing the content of ZIP files. This FUSE filesystem texpands ZIP files like folders

Topics

Resources

Stars

1 star

Watchers

1 watching

Forks

Releases

Contributors

Languages