restic

Author	SHA1	Message	Date
Michael Eischer	3424088274	Merge pull request #4644 from MichaelEischer/refactor-repair-packs Refactor and test `repair packs`	2024-01-27 13:00:51 +01:00
Michael Eischer	f0e1ad2285	fix linter warning	2024-01-27 12:51:45 +01:00
Michael Eischer	fd579421dd	repository: deduplicate test	2024-01-27 12:51:45 +01:00
Michael Eischer	42c9318b9c	repair pack: add tests	2024-01-27 12:51:45 +01:00
Michael Eischer	764b0bacd6	repair pack: add support for truncated files	2024-01-27 12:51:45 +01:00
Michael Eischer	7c351bc53c	repair pack: reenable auto index updates The method is not available on the restic.Repository interface that is used for testing. Drop the call as a small amount of additional index writes is not a problem.	2024-01-27 12:51:45 +01:00
Michael Eischer	feeab84204	repair pack: extract the repair logic into the repository package Currently, the cmd/restic package contains a significant amount of code that modifies repository internals. This code should in the mid-term move into the repository package.	2024-01-27 12:51:45 +01:00
Michael Eischer	cb50832d50	index: let MasterIndex.Save also delete obsolete indexes	2024-01-27 12:51:08 +01:00
Michael Eischer	c13bf0b607	repository: Introduce RemoveKey function This replaces directly removing keys via the backend.	2024-01-27 12:42:58 +01:00
Michael Eischer	2c310a526e	repository: Replace StreamPack function with LoadBlobsFromPack method LoadBlobsFromPack is now part of the repository struct. This ensures that users of that method don't have to deal will internals of the repository implementation. The filerestorer tests now also contain far fewer pack file implementation details.	2024-01-19 21:40:43 +01:00
Michael Eischer	6b7b5c89e9	repository: prepare StreamPack refactor	2024-01-19 21:40:43 +01:00
Michael Eischer	fb422497af	repository: split StreamPack implementation Move the actual decoding of the pack data into a separate iterator.	2024-01-19 21:39:55 +01:00
Michael Eischer	77b1c52673	repository: test that StreamPack only delivers blobs once	2024-01-07 10:54:53 +01:00
Michael Eischer	fe5c337ca2	repository: StreamPack delivers blobs at most once If an error occurred while streaming a pack file, this could result in passing some of the blobs multiple times to the callback function. This significantly complicates using StreamPack correctly and is unnecessary. Retries do not change the content of a blob and thus only deliver the same result over and over again.	2024-01-07 10:54:49 +01:00
Andrea Gelmini	241916d55b	Fix typos	2023-12-06 13:11:55 +01:00
Michael Eischer	45962c2847	Merge pull request #4499 from MichaelEischer/modular-backend-code Split backend code from restic package	2023-10-27 20:19:20 +02:00
Leo R. Lundgren	aafb806a8c	doc: Correct two typos	2023-10-27 18:56:32 +02:00
Michael Eischer	c7b770eb1f	convert MemorizeList to be repository based Ideally, code that uses a repository shouldn't directly interact with the underlying backend. Thus, move MemorizeList one layer up.	2023-10-25 23:01:35 +02:00
Michael Eischer	1b8a67fe76	move Backend interface to backend package	2023-10-25 23:00:18 +02:00
Michael Eischer	b6d79bdf6f	restic: decouple restic.Handle	2023-10-25 22:54:07 +02:00
Michael Eischer	cb9cbe55d9	repository: store oversized blobs in separate pack files Store oversized blobs in separate pack files as the blobs is large enough to warrant its own pack file. This simplifies the garbage collection of such blobs and keeps the cache smaller, as oversize (tree) blobs only have to be downloaded if they are actually used.	2023-10-17 22:52:16 +02:00
Michael Eischer	3fd0ad7448	repository: list index files only once	2023-10-01 19:53:26 +02:00
arjunajesh	ed65a7dbca	implement progress bar for index loading	2023-10-01 19:52:59 +02:00
Michael Eischer	191c47d30e	Merge pull request #4353 from MichaelEischer/tune-gc Tune Go garbage collector	2023-06-16 23:24:39 +02:00
Michael Eischer	eef0ee7a85	repository: trigger GC after loading the index Loading the index requires some scratch space, thus make sure that this memory does not factor into the targeted gc memory usage limit.	2023-06-02 21:56:14 +02:00
Michael Eischer	ffca602315	repository: Fix panic in benchmarkLoadIndex	2023-05-28 23:55:47 +02:00
Michael Eischer	d1a5ec7839	Rename unused testing parameter to _ The parameter is an additional marker that the test helper must only be used for tests.	2023-05-18 21:17:53 +02:00
Michael Eischer	1514593f22	Remove unused context or testing parameters	2023-05-18 21:17:53 +02:00
Michael Eischer	e01baeabba	Use either test or rtest to refer to internal test helpers A single test file should not use both names.	2023-05-18 21:15:45 +02:00
Michael Eischer	5773b86d02	repository: Push all usage of errors.Fatal out of the package As the `Fatal` error type only includes a string, it becomes impossible to inspect the contained error. This is for a example a problem for the fuse implementation, which must be able to detect context.Canceled errors. Co-authored-by: greatroar <61184462+greatroar@users.noreply.github.com>	2023-05-18 17:27:41 +02:00
greatroar	d129baba7a	repository: Reuse buffers in Repository.LoadUnpacked This method had a buffer argument, but that was nil at all call sites. That's removed, and instead LoadUnpacked now reuses whatever it allocates inside its retry loop.	2023-01-30 22:01:01 +01:00
Michael Eischer	1adf28a2b5	repository: properly return invalid data error in LoadUnpacked The retry backend does not return the original error, if its execution is interrupted by canceling the context. Thus, we have to manually ensure that the invalid data error gets returned. Additionally, use the retry backend for some of the repository tests, as this is the configuration which will be used by restic.	2023-01-14 17:57:02 +01:00
Michael Eischer	6d9675c323	repository: cleanup error message on invalid data The retry printed the filename twice: ``` Load(<lock/04804cba82>, 0, 0) returned error, retrying after 720.254544ms: load(<lock/04804cba82>): invalid data returned ``` now the warning has changed to ``` Load(<lock/04804cba82>, 0, 0) returned error, retrying after 720.254544ms: invalid data returned ```	2023-01-14 17:57:02 +01:00
Michael Eischer	90fb6f70b4	Merge pull request #4089 from greatroar/errors Clean up error handling further	2022-12-24 10:41:56 +01:00
greatroar	b150dd0235	all: Replace some errors.Wrap calls by errors.WithStack Mostly changed the ones that repeat the name of a system call, which is already contained in os.PathError.Op. internal/fs.Reader had to be changed to actually return such errors.	2022-12-17 09:41:07 +01:00
greatroar	c0b5ec55ab	repository: Remove empty cleanup functions in tests TestRepository and its variants always returned no-op cleanup functions. If they ever do need to do cleanup, using testing.T.Cleanup is easier than passing these functions around.	2022-12-11 11:06:25 +01:00
Michael Eischer	40ac678252	backend: remove Test method The Test method was only used in exactly one place, namely when trying to create a new repository it was used to check whether a config file already exists. Use a combination of Stat() and IsNotExist() instead.	2022-12-03 11:28:10 +01:00
Michael Eischer	ff7ef5007e	Replace most usages of ioutil with the underlying function The ioutil functions are deprecated since Go 1.17 and only wrap another library function. Thus directly call the underlying function. This commit only mechanically replaces the function calls.	2022-12-02 19:36:43 +01:00
Michael Eischer	a1eb923876	remove no longer necessary conditional compiles	2022-11-27 13:18:44 +01:00
Alexander Neumann	8dd95b710e	Merge pull request #3992 from MichaelEischer/err-on-invalid-compression Return error if RESTIC_COMPRESSION env variable is invalid	2022-11-04 19:41:34 +01:00
greatroar	137f0bc944	repository: Fix benchmarkSaveAndEncrypt	2022-10-29 23:09:17 +02:00
Michael Eischer	01f0db4e56	return error if RESTIC_COMPRESSION env variable is invalid	2022-10-29 22:03:39 +02:00
Michael Eischer	c4fc5c97f9	prune: Use a single CountedBlobSet to track blobs The set covers necessary, existing and duplicate blobs. This removes the duplicate sets used to track whether all necessary blobs also exist. This reduces the memory usage of prune by about 20-30%.	2022-10-22 18:45:12 +02:00
Michael Eischer	8d62a7adb4	identify keys by ID and not name	2022-10-15 16:07:43 +02:00
Michael Eischer	02634dce7a	restic: change Find to return ids That way consumers no longer have to manually convert the returned name to an id.	2022-10-15 16:06:54 +02:00
Michael Eischer	2e3f1c08c5	repository: split index into a separate package	2022-10-08 21:15:34 +02:00
Michael Eischer	5760ba6989	Merge pull request #3949 from MichaelEischer/simplify-mixedpacks repository: remove IsMixedPack and add replacement for checker	2022-10-08 21:14:14 +02:00
Michael Eischer	4bb5240720	repository: remove unused PrefixLength	2022-10-03 12:15:53 +02:00
Michael Eischer	999fe29976	repository: hide prepareCache	2022-10-03 12:15:53 +02:00
Michael Eischer	ddcf549eba	repository: remove IsMixedPack and add replacement for checker Repositories with mixed packs are probably quite rare by now. When loading data blobs from a mixed pack file, this will no longer trigger caching that file. However, usually tree blobs are accessed first such that this shouldn't make much of a difference. The checker gets a simpler replacement.	2022-10-03 12:03:59 +02:00
Michael Eischer	5c6b6edefe	retry index, lock and snapshot loading on hash mismatch	2022-09-25 11:35:35 +02:00
Michael Eischer	78d2312ee9	Merge pull request #3854 from MichaelEischer/sparsefiles restore: Add support for sparse files	2022-09-24 22:04:02 +02:00
Michael Eischer	c147422ba5	repository: special case SaveBlob for all zero chunks Sparse files contain large regions containing only zero bytes. Checking that a blob only contains zeros is possible with over 100GB/s for modern x86 CPUs. Calculating sha256 hashes is only possible with 500MB/s (or 2GB/s using hardware acceleration). Thus we can speed up the hash calculation for all zero blobs (which always have length chunker.MinSize) by checking for zero bytes and then using the precomputed hash. The all zeros check is only performed for blobs with the minimal chunk size, and thus should add no overhead most of the time. For chunks which are not all zero but have the minimal chunks size, the overhead will be below 2% based on the above performance numbers. This allows reading sparse sections of files as fast as the kernel can return data to us. On my system using BTRFS this resulted in about 4GB/s.	2022-09-24 21:39:39 +02:00
Michael Eischer	1ebd57247a	repository: optimize MasterIndex.Each Sending data through a channel at very high frequency is extremely inefficient. Thus use simple callbacks instead of channels. > name old time/op new time/op delta > MasterIndexEach-16 6.68s ±24% 0.96s ± 2% -85.64% (p=0.008 n=5+5)	2022-09-24 12:21:59 +02:00
Michael Eischer	825b95e313	repository: add benchmark for MasterIndex.Each	2022-09-24 12:21:59 +02:00
Michael Eischer	7682149c9d	repository: cleanup copy connection count check	2022-08-28 11:40:56 +02:00
Michael Eischer	b03277ead5	repository: don't hang when copying using a single connection	2022-08-28 11:40:31 +02:00
MichaelEischer	bee15dd555	Merge pull request #3879 from MichaelEischer/mem-optimize Some random (minor) memory-allocation optimizations	2022-08-26 20:33:02 +02:00
Michael Eischer	cc4728d287	repository: Do not report ignored packs in EachByPack Ignored packs were reported as an empty pack by EachByPack. The most immediate effect of this is that the progress bar for rebuilding the index reports processing more packs than actually exist.	2022-08-21 10:38:40 +02:00
Michael Eischer	7a992fc794	repository: Reduce buffer reallocations in ForAllIndexes Previously the buffer was grown incrementally inside `repo.LoadUnpacked`. But we can do better as we already know how large the index will be. Allocate a bit more memory to increase the chance that the buffer can be reused in the future.	2022-08-19 21:13:40 +02:00
Michael Eischer	77b1980d8e	repository: MasterIndex.Packs: reduce allocations	2022-08-19 21:10:43 +02:00
Michael Eischer	6ff9517e45	repository: MasterIndex.ListPacks / Index.EachByPack allow earlier GC Allow earlier garbage collection of some of the intermediate data structures.	2022-08-19 21:06:33 +02:00
Michael Eischer	f414db987d	gofmt all files Apparently the rules for comment formatting have changed with go 1.19.	2022-08-19 19:12:26 +02:00
Michael Eischer	7266f07c87	repository: StreamPack in parts if there are too large gaps For large pack sizes we might be only interested in the first and last blob of a pack file. Thus stream a pack file in multiple parts if the gaps between requested blobs grow too large.	2022-08-05 23:48:36 +02:00
Michael Eischer	1b076cda97	rename option to --pack-size	2022-08-05 23:47:43 +02:00
Kyle Brennan	1e3f05c3f1	repository: prevent header overfill	2022-08-05 23:47:12 +02:00
Michael Eischer	0a6fa602c8	add option for setting min pack size	2022-08-05 23:47:12 +02:00
Michael Eischer	73053674d9	repository: Test fallback to existing blobs	2022-07-30 17:37:07 +02:00
Michael Eischer	623770eebb	repository: try to recover from invalid blob while repacking If a blob that should be kept is invalid, Repack will now try to request the blob using LoadBlob. Only return an error if that fails.	2022-07-30 17:37:07 +02:00
MichaelEischer	443cc49afd	Merge pull request #3830 from MichaelEischer/cleanup-repo Extract Load/SaveTree/JSONUnpacked from repository	2022-07-23 10:46:13 +02:00
Michael Eischer	9729e6d7ef	backend: extract readerat from restic package	2022-07-17 15:29:09 +02:00
Michael Eischer	8c11fc3ec9	crypto: move crypto buffer helpers	2022-07-17 13:42:23 +02:00
Michael Eischer	89d3ce852b	repository: extract Load/StoreJSONUnpacked A Load/Store method for each data type is much clearer. As a result the repository no longer needs a method to load / store json.	2022-07-17 13:22:00 +02:00
Michael Eischer	fbcbd5318c	repository: extract LoadTree/SaveTree The repository has no real idea what a Tree is. So these methods never belonged there.	2022-07-17 13:11:28 +02:00
Lorenz Bausch	d6e3c7f28e	Wording: change repo to repository	2022-07-08 20:05:35 +02:00
Michael Eischer	6f53ecc1ae	adapt workers based on whether an operation is CPU or IO-bound Use runtime.GOMAXPROCS(0) as worker count for CPU-bound tasks, repo.Connections() for IO-bound task and a combination if a task can be both. Streaming packs is treated as IO-bound as adding more worker cannot provide a speedup. Typical IO-bound tasks are download / uploading / deleting files. Decoding / Encoding / Verifying are usually CPU-bound. Several tasks are a combination of both, e.g. for combined download and decode functions. In the latter case add both limits together. As the backends have their own concurrency limits restic still won't download more than repo.Connections() files in parallel, but the additional workers can decode already downloaded data in parallel.	2022-07-03 12:19:26 +02:00
Michael Eischer	753e56ee29	repository: Limit to a single pending pack file Use only a single not completed pack file to keep the number of open and active pack files low. The main change here is to defer hashing the pack file to the upload step. This prevents the pack assembly step to become a bottleneck as the only task is now to write data to the temporary pack file. The tests are cleaned up to no longer reimplement packer manager functions.	2022-07-02 22:42:34 +02:00
Michael Eischer	120ccc8754	repository: Rework blob saving to use an async pack uploader Previously, SaveAndEncrypt would assemble blobs into packs and either return immediately if the pack is not yet full or upload the pack file otherwise. The upload will block the current goroutine until it finishes. Now, the upload is done using separate goroutines. This requires changes to the error handling. As uploads are no longer tied to a SaveAndEncrypt call, failed uploads are signaled using an errgroup. To count the uploaded amount of data, the pack header overhead is no longer returned by `packer.Finalize` but rather by `packer.HeaderOverhead`. This helper method is necessary to continue returning the pack header overhead directly to the responsible call to `repository.SaveBlob`. Without the method this would not be possible, as packs are finalized asynchronously.	2022-07-02 22:42:34 +02:00
Michael Eischer	04c23fa95d	rebuild-index: correctly rebuild index for mixed packs For mixed packs, data and tree blobs were stored in separate index entries. This results in warning from the check command and maybe other problems.	2022-07-02 19:24:02 +02:00
Michael Eischer	a6e9e08034	Account for pack header overhead at each entry This will miss the pack header crypto overhead and the length field, which only amount to a few bytes per pack file.	2022-07-02 18:55:58 +02:00
Alexander Neumann	99634c0936	Return real size from SaveBlob	2022-07-02 18:55:12 +02:00
MichaelEischer	fdc53a9d32	Merge pull request #3787 from MichaelEischer/refactor-repository repository: (Mostly) index-related cleanups	2022-07-02 18:54:04 +02:00
Michael Eischer	ec7c9ce88b	drop unused repository.Loader interface	2022-07-02 18:39:59 +02:00
Michael Eischer	2cd7e90ad1	repository: cleanup	2022-07-02 18:39:59 +02:00
Michael Eischer	c1a8fa4290	repository: remove unused packIDToIndex field	2022-07-02 18:39:59 +02:00
Michael Eischer	e68c3a4e62	repository: simplify CreateIndexFromPacks	2022-07-02 18:39:59 +02:00
Michael Eischer	1974ad7ce2	repository: hide MasterIndex.FinalizeFullIndexes / FinalizeNotFinalIndexes	2022-07-02 18:39:59 +02:00
Michael Eischer	ef53ca4a5a	repository: remove MasterIndex.All()	2022-07-02 18:39:59 +02:00
Michael Eischer	bf81bf0795	repository: Properly set id for finalized index As MergeFinalIndex and index uploads can occur concurrently, it is necessary for MergeFinalIndex to check whether the IDs for an index were already set before merging it. Otherwise, we'd loose the ID of an index which is set _after_ uploading it.	2022-07-02 18:39:59 +02:00
Michael Eischer	e0a7852b8b	repository: remove unused (Master)Index.Count	2022-07-02 18:39:58 +02:00
Michael Eischer	8ef2968f28	repository: remove unused index.ListPack	2022-07-02 18:39:12 +02:00
Michael Eischer	e4f20dea61	repository: inline index.encode	2022-07-02 18:39:12 +02:00
Michael Eischer	fe5a8e137a	repository: remove unused index.Store	2022-07-02 18:39:12 +02:00
Michael Eischer	628ae799ca	repository: make flushPacks private	2022-07-02 18:39:12 +02:00
Michael Eischer	ed8aa15376	repository: add Save method to MasterIndex interface	2022-07-02 18:38:56 +02:00
Michael Eischer	a77d5c4d11	repository: index saving belongs into the MasterIndex	2022-07-02 18:38:56 +02:00
MichaelEischer	2c893fe43c	Merge pull request #3798 from greatroar/errors all: Move away from pkg/errors, easy cases	2022-06-17 19:01:40 +02:00
greatroar	f92ecf13c9	all: Move away from pkg/errors, easy cases github.com/pkg/errors is no longer getting updates, because Go 1.13 went with the more flexible errors.{As,Is} function. Use those instead: errors from pkg/errors already support the Unwrap interface used by 1.13 error handling. Also: * check for io.EOF with a straight ==. That value should not be wrapped, and the chunker (whose error is checked in the cases changed) does not wrap it. * Give custom Error methods pointer receivers, so there's no ambiguity when type-switching since the value type will no longer implement error. * Make restic.ErrAlreadyLocked private, and rename it to alreadyLockedError to match the stdlib convention that error type names end in Error. * Same with rest.ErrIsNotExist => rest.notExistError. * Make s3.Backend.IsAccessDenied a private function.	2022-06-14 08:36:38 +02:00
Jayson Wang	f144920ed5	fix handling of maxKeys in SearchKey	2022-06-12 14:19:06 +02:00
greatroar	c9557b2822	internal/repository: Fix LoadBlob + fuzz test When given a buf that is big enough for a compressed blob but not its decompressed contents, the copy at the end of LoadBlob would skip the last part of the contents. Fixes #3783.	2022-06-06 17:02:28 +02:00
MichaelEischer	b2a2e5f727	Merge pull request #3753 from greatroar/indexmap-alloc repository: Re-tune indexmap allocation strategy	2022-05-14 15:44:08 +02:00
greatroar	5141228e0c	repository: Re-tune indexmap allocation strategy `fd05037e1a` changed the allocation batch size from 256 to 128 under the assumption that an indexEntry is 60 bytes on amd64, but it's 64: structs are padded out to a multiple of 8 for alignment reasons. That means we'd waste no space in malloc even without the batch allocation, at least on 64-bit machines. While that strategy cuts the overallocation down dramatically for many small indexes, it also seems to slow allocation down (Go 1.18, Linux, amd64, -benchtime=2s): name old time/op new time/op delta DecodeIndex-8 4.67s ± 5% 4.60s ± 1% ~ (p=0.953 n=10+5) DecodeIndexParallel-8 4.67s ± 3% 4.60s ± 1% ~ (p=0.953 n=10+5) IndexHasUnknown-8 37.8ns ± 8% 36.5ns ±14% ~ (p=0.841 n=5+5) IndexHasKnown-8 38.5ns ±12% 37.7ns ±10% ~ (p=0.968 n=5+5) IndexAlloc-8 615ms ±18% 607ms ± 1% ~ (p=1.000 n=10+5) IndexAllocParallel-8 245ms ±11% 285ms ± 6% +16.40% (p=0.001 n=10+5) MasterIndexAlloc-8 286ms ± 9% 275ms ± 2% ~ (p=1.000 n=10+5) LoadIndex/v1-8 27.0ms ± 4% 26.8ms ± 1% ~ (p=0.690 n=5+5) LoadIndex/v2-8 22.4ms ± 1% 22.8ms ± 2% +1.48% (p=0.016 n=5+5) name old alloc/op new alloc/op delta IndexAlloc-8 446MB ± 0% 446MB ± 0% -0.00% (p=0.000 n=8+4) IndexAllocParallel-8 446MB ± 0% 446MB ± 0% -0.00% (p=0.008 n=8+5) MasterIndexAlloc-8 213MB ± 0% 159MB ± 0% -25.47% (p=0.000 n=10+5) name old allocs/op new allocs/op delta IndexAlloc-8 913k ± 0% 2632k ± 0% +188.19% (p=0.008 n=5+5) IndexAllocParallel-8 913k ± 0% 2632k ± 0% +188.21% (p=0.008 n=5+5) MasterIndexAlloc-8 318k ± 0% 1172k ± 0% +267.86% (p=0.008 n=5+5) Instead, this patch sets a batch size of 4, which means no space is wasted by malloc on 64-bit and very little on 32-bit. It still gets very close to the savings from not allocating in batches, without requiring special code for bits.UintSize==64. Benchmark results, again for Linux/amd64: name old time/op new time/op delta DecodeIndex-8 4.67s ± 5% 4.83s ± 9% ~ (p=0.315 n=10+10) DecodeIndexParallel-8 4.67s ± 3% 4.68s ± 4% ~ (p=0.315 n=10+10) IndexHasUnknown-8 37.8ns ± 8% 44.5ns ±19% ~ (p=0.095 n=5+5) IndexHasKnown-8 38.5ns ±12% 36.9ns ± 8% ~ (p=0.690 n=5+5) IndexAlloc-8 615ms ±18% 628ms ±18% ~ (p=0.218 n=10+10) IndexAllocParallel-8 245ms ±11% 262ms ± 9% +7.02% (p=0.043 n=10+10) MasterIndexAlloc-8 286ms ± 9% 287ms ±13% ~ (p=1.000 n=10+10) LoadIndex/v1-8 27.0ms ± 4% 26.8ms ± 0% ~ (p=1.000 n=5+5) LoadIndex/v2-8 22.4ms ± 1% 22.5ms ± 0% ~ (p=0.056 n=5+5) name old alloc/op new alloc/op delta IndexAlloc-8 446MB ± 0% 446MB ± 0% ~ (p=1.000 n=8+10) IndexAllocParallel-8 446MB ± 0% 446MB ± 0% -0.00% (p=0.000 n=8+8) MasterIndexAlloc-8 213MB ± 0% 160MB ± 0% -25.02% (p=0.000 n=10+9) name old allocs/op new allocs/op delta IndexAlloc-8 913k ± 0% 1333k ± 0% +45.94% (p=0.000 n=8+10) IndexAllocParallel-8 913k ± 0% 1333k ± 0% +45.94% (p=0.000 n=8+8) MasterIndexAlloc-8 318k ± 0% 525k ± 0% +64.99% (p=0.000 n=10+10) The allocation method indexmap.newEntry has also been rewritten in a form that is a few instructions shorter.	2022-05-11 21:22:14 +02:00
MichaelEischer	df554e5f69	Merge pull request #3748 from greatroar/runworkers repository: Remove RunWorkers, report ctx.Err()	2022-05-11 19:38:46 +02:00
greatroar	2e0f1f5113	repository: Remove RunWorkers, report ctx.Err() This removes RunWorkers, which had become mere overhead by successive refactors. It also ensures that each former user of that function returns any context error that occurs, so failure to complete an operation is always reported as an error.	2022-05-10 22:26:00 +02:00
Michael Eischer	ae7e51382a	Fix error on temp file deletion on windows Apparently it can take a moment between closing a tempfile marked as DELETE_ON_CLOSE and it actually being deleted. During that time the file is inaccessible. Thus just skip deleting the temp file on windows.	2022-05-09 22:43:26 +02:00
Michael Eischer	cf5cb673fb	repository: Use existing method to collect pack ids	2022-04-30 19:14:21 +02:00
Michael Eischer	b335cb6285	repository: Refactor index IDs collection	2022-04-30 19:14:21 +02:00
Michael Eischer	4b01b06f2f	repository: Test compressed blobs in StreamPack	2022-04-30 11:34:10 +02:00
Michael Eischer	ec2b25565a	repository: test uncompressedLength field and index example	2022-04-30 11:34:10 +02:00
Michael Eischer	9ffb8920f1	repository: run blackbox tests using old and new repo version	2022-04-30 11:34:10 +02:00
Michael Eischer	abe5935693	repository: unify repository version-specific initialization Mark the master index as compressed also when initializing a new repository. This is only relevant for testing.	2022-04-30 11:34:10 +02:00
Alexander Neumann	8776031f96	Leave allocating slices to the decompress code	2022-04-30 11:34:10 +02:00
Alexander Neumann	5eb05a0afe	Configure zstd encoder/decoder	2022-04-30 11:34:10 +02:00
Michael Eischer	2f36e044db	Cleanup pack header check	2022-04-30 11:34:10 +02:00
Alexander Neumann	8b11b86383	Add option global --compression	2022-04-30 11:34:10 +02:00
Michael Eischer	7132df529e	repository: Increase index size for repo version 2 A compressed index is only about one third the size of an uncompressed one. Thus increase the number of entries in an index to avoid cluttering the repository with small indexes.	2022-04-30 11:34:10 +02:00
Michael Eischer	66f9048bce	repository: Alloc zstd encoder/decoder on demand	2022-04-30 11:34:10 +02:00
Michael Eischer	fd05037e1a	repository: recalibrate index batch allocation size	2022-04-30 11:34:10 +02:00
Michael Eischer	6fb408d90e	repository: implement pack compression	2022-04-30 11:34:10 +02:00
Michael Eischer	362ab06023	init: Add flag to specify created repository version	2022-04-30 10:07:42 +02:00
Michael Eischer	4b957e7373	repository: Implement index/snapshot/lock compression The config file is not compressed as it should remain readable by older restic versions such that these can return a proper error. As the old format for unpacked data does not include a version header, make use of a trick: The old data is always encoded as JSON. Thus it can only start with '{' or '['. For any other value the first byte indicates a versioned format. The version is set to 2 for now. Then the zstd compressed data follows.	2022-04-30 10:07:42 +02:00
Michael Eischer	e597b99b55	repository: Reduce repack workers to prevent deadlock As repack streams packs these occupy one backend connection. Uploading a new pack also requires a backend connection. To prevent a deadlock during repack when reaching the backend connections limit, simply limit the repackWorker count to always leave one connection for uploading.	2022-04-23 11:28:18 +02:00
Alexander Neumann	a059ef90f8	Merge pull request #3702 from MichaelEischer/extend-config-error Print used key name if config fails to load	2022-04-10 20:25:24 +02:00
Michael Eischer	c2aabb2686	Print used key name if config fails to load	2022-04-09 22:38:18 +02:00
Alexander Neumann	04e054465a	Merge pull request #3475 from MichaelEischer/local-sftp-conn-limit Limit concurrent operations for local / sftp backend	2022-04-09 21:33:00 +02:00
Michael Eischer	7b9ae91e04	copy: Load snapshots before indexes	2022-04-09 12:27:25 +02:00
Michael Eischer	cd783358d3	local: Limit concurrent backend operations Use a limit of 2 similar to the filereader concurrency in the archiver.	2022-04-09 12:21:38 +02:00
Michael Eischer	6408686973	repository: Simplify Blob equality check	2022-03-28 22:09:49 +02:00
Michael Eischer	243698680a	crypto: Use helpers for size calculations	2022-03-28 22:09:49 +02:00
Michael Eischer	f78bd14e28	repository: Remove pack implementation details from MasterIndex	2022-03-28 22:09:49 +02:00
Michael Eischer	dc3d77dacc	repository: make saveAndEncrypt private	2022-03-28 22:09:49 +02:00
Michael Eischer	6877e7edbb	repository: Rename LoadAndDecrypt to LoadUnpacked The method is the complement for SaveUnpacked and not for SaveAndEncrypt. The latter assembles blobs into pack files.	2022-03-28 22:09:49 +02:00
Michael Eischer	537b4c310a	copy: Implement by reusing repack The repack operation copies all selected blobs from a set of pack files into new pack files. For prune the source and destination repositories are identical. To implement copy, just use a different source and destination repository.	2022-03-26 20:47:15 +01:00
Alexander Neumann	e682f7c0d6	Add tests for StreamPack	2022-03-21 21:15:03 +01:00
Michael Eischer	bba8ba7a5b	repository: cancel streampack context after error	2022-02-12 20:18:25 +01:00
Michael Eischer	47554a3428	repository: Fix error handling in repack When storing a blob fails, this is a fatal error which must not be retried.	2022-02-12 20:18:25 +01:00
Michael Eischer	930a00ad54	checker: reuse bufio reader	2022-02-12 20:18:25 +01:00
Michael Eischer	34ebafb8b6	repository: don't crash if blob size is too short	2022-02-12 20:18:25 +01:00
Michael Eischer	becebf5d88	repository: remove unused DownloadAndHash	2022-02-12 20:18:25 +01:00
Michael Eischer	f1e58e7c7f	checker: rewrite ReadData to stream packs	2022-02-12 20:18:25 +01:00
Michael Eischer	f40abd92fa	restorer: convert to use StreamPack	2022-02-12 20:18:25 +01:00
Michael Eischer	f00f690658	repository: stream packs during repacking	2022-02-12 20:18:25 +01:00
Michael Eischer	c4a2bfcb39	repository: Add StreamPacks function The function supports efficiently loading a specified list of blobs from a single pack in a streaming fashion. That is there's no need for temporary files independent of the pack size.	2022-02-12 20:18:25 +01:00
Michael Eischer	153e2ba859	repository: Implement lisiting blobs per pack file	2022-02-12 20:18:24 +01:00
greatroar	8d2996eaaa	Replace siphash by hash/maphash In Go 1.17.1, maphash has become quite a bit faster than siphash, so we can drop one third-party dependency. maphash is just an interface to the standard Go map's hash function, which we already trust for other use cases. Benchmark results on linux/amd64, -benchtime=3s: name old time/op new time/op delta IndexHasUnknown-8 50.6ns ±10% 41.0ns ±19% -18.92% (p=0.000 n=9+10) IndexHasKnown-8 52.6ns ±12% 41.5ns ±12% -21.13% (p=0.000 n=9+10) IndexMapHash-8 3.64µs ± 1% 2.00µs ± 0% -45.09% (p=0.000 n=10+9) IndexAlloc-8 700ms ± 1% 601ms ± 6% -14.18% (p=0.000 n=8+10) IndexAllocParallel-8 205ms ± 5% 192ms ± 8% -6.18% (p=0.043 n=10+10) MasterIndexAlloc-8 319ms ± 1% 279ms ± 5% -12.58% (p=0.000 n=10+10) MasterIndexLookupSingleIndex-8 156ns ± 8% 147ns ± 6% -5.46% (p=0.023 n=10+10) MasterIndexLookupMultipleIndex-8 150ns ± 7% 142ns ± 8% -5.69% (p=0.007 n=10+10) MasterIndexLookupSingleIndexUnknown-8 74.4ns ± 6% 72.0ns ± 9% ~ (p=0.175 n=10+9) MasterIndexLookupMultipleIndexUnknown-8 67.4ns ± 9% 65.5ns ± 7% ~ (p=0.340 n=9+9) MasterIndexLookupParallel/known,indices=25-8 461ns ± 2% 445ns ± 2% -3.49% (p=0.000 n=10+10) MasterIndexLookupParallel/unknown,indices=25-8 408ns ±11% 378ns ± 5% -7.22% (p=0.035 n=10+9) MasterIndexLookupParallel/known,indices=50-8 479ns ± 1% 437ns ± 4% -8.82% (p=0.000 n=10+10) MasterIndexLookupParallel/unknown,indices=50-8 406ns ± 8% 343ns ±15% -15.44% (p=0.001 n=10+10) MasterIndexLookupParallel/known,indices=100-8 480ns ± 1% 455ns ± 5% -5.15% (p=0.000 n=8+10) MasterIndexLookupParallel/unknown,indices=100-8 391ns ±18% 382ns ± 8% ~ (p=0.315 n=10+10) MasterIndexLookupBlobSize-8 71.0ns ± 8% 57.2ns ±11% -19.36% (p=0.000 n=9+10) PackerManager-8 254ms ± 1% 254ms ± 1% ~ (p=0.285 n=15+15) name old speed new speed delta IndexMapHash-8 1.12GB/s ± 1% 2.05GB/s ± 0% +82.13% (p=0.000 n=10+9) PackerManager-8 208MB/s ± 1% 207MB/s ± 1% ~ (p=0.281 n=15+15) name old alloc/op new alloc/op delta IndexMapHash-8 0.00B 0.00B ~ (all equal) IndexAlloc-8 400MB ± 0% 400MB ± 0% ~ (p=1.000 n=9+10) IndexAllocParallel-8 401MB ± 0% 401MB ± 0% +0.00% (p=0.000 n=10+10) MasterIndexAlloc-8 258MB ± 0% 262MB ± 0% +1.42% (p=0.000 n=9+10) PackerManager-8 73.1kB ± 0% 73.1kB ± 0% ~ (p=0.382 n=13+13) name old allocs/op new allocs/op delta IndexMapHash-8 0.00 0.00 ~ (all equal) IndexAlloc-8 907k ± 0% 907k ± 0% -0.00% (p=0.000 n=10+10) IndexAllocParallel-8 907k ± 0% 907k ± 0% +0.00% (p=0.009 n=10+10) MasterIndexAlloc-8 327k ± 0% 317k ± 0% -3.06% (p=0.000 n=10+10) PackerManager-8 744 ± 0% 744 ± 0% ~ (all equal)	2021-09-19 16:05:18 +02:00
Alexander Weiss	81876d5c1b	Simplify cache logic	2021-09-03 21:01:00 +02:00
Michael Eischer	9aa2eff384	Add plumbing to calculate backend specific file hash for upload This enables the backends to request the calculation of a backend-specific hash. For the currently supported backends this will always be MD5. The hash calculation happens as early as possible, for pack files this is during assembly of the pack file. That way the hash would even capture corruptions of the temporary pack file on disk.	2021-08-04 22:17:46 +02:00
Ryan Hitchman	77bf148460	backup: add --dry-run/-n flag to show what would happen. This can be used to check how large a backup is or validate exclusions. It does not actually write any data to the underlying backend. This is implemented as a simple overlay backend that accepts writes without forwarding them, passes through reads, and generally does the minimal necessary to pretend that progress is actually happening. Fixes #1542 Example usage: $ restic -vv --dry-run . \| grep add new /changelog/unreleased/issue-1542, saved in 0.000s (350 B added) modified /cmd/restic/cmd_backup.go, saved in 0.000s (16.543 KiB added) modified /cmd/restic/global.go, saved in 0.000s (0 B added) new /internal/backend/dry/dry_backend_test.go, saved in 0.000s (3.866 KiB added) new /internal/backend/dry/dry_backend.go, saved in 0.000s (3.744 KiB added) modified /internal/backend/test/tests.go, saved in 0.000s (0 B added) modified /internal/repository/repository.go, saved in 0.000s (20.707 KiB added) modified /internal/ui/backup.go, saved in 0.000s (9.110 KiB added) modified /internal/ui/jsonstatus/status.go, saved in 0.001s (11.055 KiB added) modified /restic, saved in 0.131s (25.542 MiB added) Would add to the repo: 25.892 MiB	2021-08-04 21:19:29 +02:00
Alexander Neumann	aef3658a5f	Address review comments	2021-01-30 20:02:37 +01:00
Alexander Neumann	16313bfcc9	errcheck: Add error check for MergeFinalIndexes()	2021-01-30 20:02:37 +01:00

1 2 3 4 5 ...

420 commits