Skip to content

git_index_add ignores file/dir conflicts #7160

Description

@Batchyx

The problem

git_index_add is documented to:

Add or update an index entry from an in-memory struct

If a previous index entry exists that has the same path and stage as the given 'source_entry', it will be replaced. Otherwise, the 'source_entry' will be added.

A full copy (including the 'path' string) of the given 'source_entry' will be inserted on the index.

One could assume from this documentation that inserting a file/dir or blob/tree conflict will fail with an error, just like creat() or mkdir() would fail on a file/dir conflict. i.e. that this test would work:

void test_index_collision__add_blob_with_conflicting_tree_wanted(void)
{
        git_index_entry entry;
        git_tree_entry *tentry;
        git_oid tree_id;
        git_tree *tree;

        memset(&entry, 0, sizeof(entry));
        entry.ctime.seconds = 12346789;
        entry.mtime.seconds = 12346789;
        entry.mode  = 0100644;
        entry.file_size = 0;
        git_oid_cpy(&entry.id, &g_empty_id);

        entry.path = "some blob";
        cl_git_pass(git_index_add(g_index, &entry));

        entry.path = "another blob";
        cl_git_pass(git_index_add(g_index, &entry));

        entry.path = "blobtree/conflict";
        cl_git_pass(git_index_add(g_index, &entry));

        /* creating a blob/tree collision should fail */
        entry.path = "blobtree";
        cl_git_fail(git_index_add(g_index, &entry));
}

What happens instead is that git_index_add does nothing and returns success:

void test_index_collision__add_blob_with_conflicting_tree_actual(void)
{
        git_index_entry entry;
        git_tree_entry *tentry;
        git_oid tree_id;
        git_tree *tree;

        memset(&entry, 0, sizeof(entry));
        entry.ctime.seconds = 12346789;
        entry.mtime.seconds = 12346789;
        entry.mode  = 0100644;
        entry.file_size = 0;
        git_oid_cpy(&entry.id, &g_empty_id);

        entry.path = "some blob";
        cl_git_pass(git_index_add(g_index, &entry));

        entry.path = "another blob";
        cl_git_pass(git_index_add(g_index, &entry));

        entry.path = "blobtree/conflict";
        cl_git_pass(git_index_add(g_index, &entry));

        /* Check blobtree/conflict exists here */
        cl_git_pass(git_index_write_tree(&tree_id, g_index));
        cl_git_pass(git_tree_lookup(&tree, g_repo, &tree_id));
        cl_git_pass(git_tree_entry_bypath(&tentry, tree, "blobtree/conflict"));
        git_tree_entry_free(tentry);
        git_tree_free(tree);

        /* Create a blob/tree collision... */
        entry.path = "blobtree";
        cl_git_pass(git_index_add(g_index, &entry));

        /* ... and nothing changed */
        cl_git_pass(git_index_write_tree(&tree_id, g_index));
        cl_git_pass(git_tree_lookup(&tree, g_repo, &tree_id));
        cl_git_pass(git_tree_entry_bypath(&tentry, tree, "blobtree/conflict"));
        git_tree_entry_free(tentry);
        git_tree_free(tree);
}

This behavior is surprising and not practical when programmatically resolving merge conflicts. If the calling code need to check for blob/tree conflicts instead of libgit2, it have to use resource-intensive workarounds since it cannot depend on internal representation details (i.e. it cannot assume that the index is sorted, so must iterate everything, or maintain its own index).

What does git do

The equivalent git command for git_index_add is git update-index --add --cacheinfo and faced with a dir/file conflict, it fails:

$ git update-index --add --cacheinnfo 100644,e69de29bb2d1d6434b8b29ae775ad8c2e48c5391,blobtree
error: 'blobtree' appears as both a file and as a directory
error: blobtree: cannot add to the index - missing --add option?
fatal: Unable to process path blobtree

Adding a file with git update-index --add blobtree also fails with the same error message.

git add behaves differently: it erases all conflicting entries before adding a file, but it is also consistent with git add deleting files/directories from the index if they do not exist on the working tree.

What actually happens

git_index_add calls index_insert(), which:

  • sorts the index
  • calls index_existing_and_best, but blobtree does not exist at any stage, so it sets existing and best to NULL and position to 0.
  • calls check_file_directory_collision with position = 0
    • calls has_file_name which returns quickly because index->entries[0] is "another blob" and does not start with blobtree.
    • calls has_dir_name which returns quickly because there is no slash in "blobtree".
  • ends up inserting the entry

But when git_index_write_tree passes over it, it iterates blobtree before blobtree/conflict, so the tree is inserted last in treebuilder.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions