treeverse / lakefs Goto Github PK

lakeFS - Data version control for your data lake | Git for data

License: Apache License 2.0

Go 75.59% Makefile 0.38% HTML 0.01% CSS 0.36% JavaScript 7.10% Dockerfile 0.10% Shell 0.24% Thrift 1.29% Scala 3.95% Python 3.60% Java 4.76% C++ 0.15% TypeScript 1.85% Lua 0.60% Ruby 0.01% Batchfile 0.02%

apache-spark apache-sparksql aws-s3 azure-blob-storage azure-storage data-engineering data-lake data-quality data-version-control data-versioning datalake datalakes git-for-data go golang google-cloud-storage hadoop-filesystem lakefs object-storage

lakefs's People

Contributors

Stargazers

Watchers

Forkers

d-harel ngaut cregev miko-code foxxnuaa deepakksahu dieptran43 jengjeng pr1-xd eylonronen naveen-kumar-r viswalahiri kzuri forkkit holajiawei souvikinator spiritedawayseasong lavanyamanohar11 ari-hacks avi111111 eshanvohra khushijindal nirupam090 aeyshubh karanprajapati8750 perry-contribs sombit-g larkceresin gauravgoyal2324 iamrishabh07 ayushassest mofolactic samhack05 coatest 0ze3r0 sachikant786 goobar07 mishrajiharsh219 strong10mede deepshi1410 sajal243 gauravsaha-97 rohansahana jbampton bhupiji kutubkhan2005 rudrakj m-prerna mathagician infinitelloop iamdopecode codeonix nishith-1997 yashs911 yashmit178 supriya-ld atievewadhwa simarpreetsingh-019 ravitejavelamuri alenros shubham1204 impramodsargar amit1173 nikita812 bhavik9988 jarrodhroberson indhupriya vibhuti1402agg sufiyan1997 rahul-160 kaushalag29 shamikakumar kishoreraju2 sarathsp06 avats-dev itaiad200 mittalkartik2000 ymrmmm sonam2905 avmi aryansharmaa anshssonkhia daniel-shuy sohamds sumindar jeffthomas2000 eshamahendra sinithh aidhamza slumbi danielleholtz ibearddev isgasho avshalomman acautman leo2904 christophergrant pranavsriram8 h2atecnologia sunilbhara

lakefs's Issues

CI step: validate dependencies comply with license

Refactor list entries by level's code

Create lakeFS Helm chart

configuration: search for config file under /etc/lakefs.yaml (for nicer docker run)

Expiry under GCS

Explain how to run "lakefs expire" periodically in a deployed system

this should be added to the Deployment documentation as well.

Originally posted by @ozkatz in https://github.com/treeverse/lakeFS/pull/230/files

Move "/setup_lakefs" endpoint to the API

It's currently a standalone endpoint outside of swagger therefore hidden.
Move that endpoint to swagger and document properly.

System tests discovery

MVCC rollback branch changes

The capability to rollback branch committed changes by apply the reverse changes.
Optionally for specific range of changes/commits

Allow reuse of db container when testing locally

In a developer environment, allow developer to keep db alive after the tests are finished. Also, allow running tests multiple times on the same db container.

lakeFS release binaries are not signed (macOS)

consider signing the binaries in order to have better download and execute experience on macOS .
https://goreleaser.com/customization/sign/

Test issuee

webui: ref component show only one reference back

performance testing of DB /cataloger

SQL error is shown to the user

When trying to create a repository with a name of an existing repository, the following error is printed:

Error executing command: error creating repository: insert repository: ERROR: duplicate key value violates unique constr aint "catalog_repositories_name_uindex" (SQLSTATE 23505)

This should be changed to a user-friendly error.

Example command to reproduce:
lakectl repo create lakefs://existing-repo s3://example-bucket

Metadata retention dicovery

Postponed until S3 based retention is in-place.

automate: latest release README.md to hub.docker

Deployment from commit/branch on AWS discovery

Client/server upgrade story

At the very least:

Detect errors when client and server Swagger definitions don't match for a query
Decide for both FC and BC
Can initially be as simple as allowing extra props on client parsing but disallowing them all on server parsing

Possibly required for release: these clients or servers will be involved in users' next upgrades.

Monitor repo retention

I would suggest tracking skipped repos due to errors - if one of them failed, return a non-zero return code. Otherwise it'll get logged into the void and probably won't ever be noticed by anyone until the S3 bills start racking up...

Originally posted by @ozkatz in #309 (comment)

Use testing short flag to skip playback tests

Using go test -short will skip the gateway playback testing.
It will enable faster basic test run without download and play of playback data.

lakefs generate config

In the UI, make it easily possible to copy the path to an object

In S3 web interface, the path is part of the URL so the slash separator is preserved, making the path easy to copy (see photo).

In our UI, the path is given as a query param, so slashes are escaped.
Perhaps add a copy button for the path. Need to decide which path type to copy - s3 or lakeFS, and whether it will include the repo and branch names.

commit not ordered by time

In the web UI in the commits tab
the commits are not sorted by time

lakeFS running with GCS using S3 storage interface

NO SUPPORT FOR

multipart upload
expiry

Serve swagger-ui from CDN

We include a full copy of swagger-ui in docs/assets/js. This increases size and reduces performance.

docs: remove pre-alpha warning

Architecture page in the docs currently shows at pre-alpha warning which is no longer true

Refactor list entries by level code

Repository creation first commit should include the user/author that performed the action

Sometimes the lakeFS system performs commits as a result of another user action. Examples include branch creation, repository creation, merge commits and import API commits.

In the case of branch and repository creation, an empty committer name is used (look for the constant CatalogerCommitter). This should be changed to the user that initiated the action.

The lakeFS CLA link leads no where....

On the community page. Click it, and you'll see :-)

GCS block adapter (without expire)

Improve "lakefs diagnose" to verify AWS IAM role used in retention batch tagging operation

After #405, There are numerous "interesting" failure modes for the role used for retention batch tagging:

Must be assume-able by batch tagging.
Must have put-object-tagging permission.
Must be allowed to list-bucket.
This one is odd; here's why: If not, then when an object does not exist we receive AccessDenied instead of NoSuchKey. That makes reports useless both for users and for (future) object_dedup cleanups.

Diagnose and report all of these (assuming of course that the running user has sufficient permissions...).

Fix serialize-javascript vulnerability found in webui/package-lock.json

Error when creating repo with ':' in path

SignatureDoesNotMatch due to unicode encoding in signature.

Allow multipart CopyObject (incl. with range)

System tests skeleton and basic test

Create the test binary with a simple single test.
Add github workflow to run it with every PR.

Interactive tool for generating a lakeFS configuration file

The lakefs binary uses a yaml configuration file.

As a lakeFS user, I would like a command line utility to generate this config file interactively.

Upon running "lakefs config", the user should be asked to fill in the basic information for a lakefs server to run.
A minimal configuration file example can be found here in the docs.

Example of information to get from the user:

Database connection string
Blockstore type.
If the blockstore type is S3, choose the region and authentication method (aws profile, aws key pair, or using the instance role)
S3 gateway domain name

The output should be a yaml file containing the configuration. The user should be able to specify an output destination for the file. If not specified, it should be saved to the default location: $HOME/.lakefs.yaml (with override protection).

Go linter with basic ruleset on CI for new/modified code

BA: basic configuration and makefile that uses golangci-lint was added

sig v2 match basedomain request port

sig v2 doesn't ignore port on request when compare to baredomain.
not it also panic - which should be removed or replace with error

MVCC sql builder should not concat values (gosec)

Fix printing in "lakefs auth policies"

Detected in #256. Policies are printed with lots of %v, which ends up throwing addresses at the user.

E.g. policies create:

ariels@ariels:~/Dev/lakeFS$ echo '{"id": "foo", "statement": [{"action": ["auth:*"], "effect": "Allow", "resource": "arn:::::"}]}' | ./lakectl auth policies create --policy-document -
Policy created successfully.
ID: 0xc00043c020
Creation Date: 2020-07-08 12:16:26 +0300 IDT
Statements:

+--------------+-------------------------------+-------------+--------------+--------------+---------+
| POLICY ID    | CREATION DATE                 | STATEMENT # | RESOURCE     | EFFECT       | ACTIONS |
+--------------+-------------------------------+-------------+--------------+--------------+---------+
| 0xc0004d2590 | 1970-01-01 02:00:00 +0200 IST |           0 | 0xc00043c040 | 0xc00043c030 | auth:*  |
+--------------+-------------------------------+-------------+--------------+--------------+---------+

E.g. policies list:

ariels@ariels:~/Dev/lakeFS$ ./lakectl auth policies list
+--------------+-------------------------------+-------------+--------------+--------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
| POLICY ID    | CREATION DATE                 | STATEMENT # | RESOURCE     | EFFECT       | ACTIONS                                                                                                                                                                                                                   |
+--------------+-------------------------------+-------------+--------------+--------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+
| 0xc0004e3d60 | 2020-07-07 18:28:24 +0300 IDT |           0 | 0xc0004e3d80 | 0xc0004e3d70 | auth:*                                                                                                                                                                                                                    |
| 0xc0004e3d90 | 2020-07-07 18:28:24 +0300 IDT |           0 | 0xc0004e3dc0 | 0xc0004e3db0 | auth:CreateCredentials, auth:DeleteCredentials, auth:ListCredentials, auth:ReadCredentials                                                                                                                                |
| 0xc0004e3dd0 | 2020-07-07 18:28:24 +0300 IDT |           0 | 0xc0004e3df0 | 0xc0004e3de0 | fs:*                                                                                                                                                                                                                      |
| 0xc0004e3e00 | 2020-07-07 18:28:24 +0300 IDT |           0 | 0xc0004e3e20 | 0xc0004e3e10 | fs:List*, fs:Read*                                                                                                                                                                                                        |
| 0xc0004e3e30 | 2020-07-07 18:28:24 +0300 IDT |           0 | 0xc0004e3e50 | 0xc0004e3e40 | fs:ListRepositories, fs:ReadRepository, fs:ReadCommit, fs:ListBranches, fs:ListObjects, fs:ReadObject, fs:WriteObject, fs:DeleteObject, fs:RevertBranch, fs:ReadBranch, fs:CreateBranch, fs:DeleteBranch, fs:CreateCommit |
| 0xc0004e3e60 | 2020-07-07 18:28:24 +0300 IDT |           0 | 0xc0004e3e80 | 0xc0004e3e70 | retention:*                                                                                                                                                                                                               |
| 0xc0004e3e90 | 2020-07-07 18:28:24 +0300 IDT |           0 | 0xc0004e3eb0 | 0xc0004e3ea0 | retention:Get*                                                                                                                                                                                                            |
| 0xc0004e3ec0 | 2020-07-08 12:16:26 +0300 IDT |           0 | 0xc0004e3ee0 | 0xc0004e3ed0 | auth:*                                                                                                                                                                                                                    |
+--------------+-------------------------------+-------------+--------------+--------------+---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------+

Playground with notebooks and public access discovery

Split "Deployment" section of docs

Right now it's a huge doc - split it to steps.

Also add Kubernetes as an installation method.

Rename "lakefs init" to "lakefs setup"

The command lakefs init should be renamed to lakefs setup, to keep consistency with the API.

Improve server response for uncommitted changes

DiffUncommitted should support prefix. It will enable us getting diff by level.
The response should include flag saying if there was a change in the branch level and under the prefix return the diff by level.
It should consider add/delete in the folder level.