How to produce a permutation of a Dask DataFrame

up vote
0
down vote

favorite

I know this topic is fairly discussed but it's still not completely clear to me what the most standard way to produce a permutation of a Dask DataFrame is.

To produce a random index in a non distributed fashion the first thing someone would likely try would be

df['random_index'] = np.random.permutation(len(df))

But in the context of Dask len(df) will trigger a computation. It's not clear to me whether invoking such a computation to realise the length makes sense. An alternative I see instead is to do something like

ds = ds.map(

    lambda (col_1, col_2): (

         <random_string>, col_1, col_2

    )

)

this will create lazily a new pseudorandom column that can be used as in index. Do you see anything wrong with that? I guess to create the random strings someone should use a good hashing algorithm to make sure the keys are evenly distributed. I was thinking of something like that

import hashlib

from random import randint  



hashlib.sha1(bytes(randint(1, 1e16))).hexdigest()

that is both fairly fast and will produce evenly distributed keys. Let me know if I am falling into any obvious pitfall(?)

edit

actually there is no reason to use a string instead of the plain integer. you just need to make sure you produce much bigger indexes than the size of your dataset

asked Nov 19 at 13:20

LetsPlayYahtzee

1,89222036

add a comment |

up vote
0
down vote

favorite

I know this topic is fairly discussed but it's still not completely clear to me what the most standard way to produce a permutation of a Dask DataFrame is.

To produce a random index in a non distributed fashion the first thing someone would likely try would be

df['random_index'] = np.random.permutation(len(df))

ds = ds.map(

    lambda (col_1, col_2): (

         <random_string>, col_1, col_2

    )

)

import hashlib

from random import randint  



hashlib.sha1(bytes(randint(1, 1e16))).hexdigest()

that is both fairly fast and will produce evenly distributed keys. Let me know if I am falling into any obvious pitfall(?)

edit

actually there is no reason to use a string instead of the plain integer. you just need to make sure you produce much bigger indexes than the size of your dataset

asked Nov 19 at 13:20

LetsPlayYahtzee

1,89222036

add a comment |

up vote
0
down vote

favorite

I know this topic is fairly discussed but it's still not completely clear to me what the most standard way to produce a permutation of a Dask DataFrame is.

To produce a random index in a non distributed fashion the first thing someone would likely try would be

df['random_index'] = np.random.permutation(len(df))

ds = ds.map(

    lambda (col_1, col_2): (

         <random_string>, col_1, col_2

    )

)

import hashlib

from random import randint  



hashlib.sha1(bytes(randint(1, 1e16))).hexdigest()

that is both fairly fast and will produce evenly distributed keys. Let me know if I am falling into any obvious pitfall(?)

edit

actually there is no reason to use a string instead of the plain integer. you just need to make sure you produce much bigger indexes than the size of your dataset

asked Nov 19 at 13:20

LetsPlayYahtzee

1,89222036

I know this topic is fairly discussed but it's still not completely clear to me what the most standard way to produce a permutation of a Dask DataFrame is.

To produce a random index in a non distributed fashion the first thing someone would likely try would be

df['random_index'] = np.random.permutation(len(df))

ds = ds.map(

    lambda (col_1, col_2): (

         <random_string>, col_1, col_2

    )

)

import hashlib

from random import randint  



hashlib.sha1(bytes(randint(1, 1e16))).hexdigest()

that is both fairly fast and will produce evenly distributed keys. Let me know if I am falling into any obvious pitfall(?)

edit

actually there is no reason to use a string instead of the plain integer. you just need to make sure you produce much bigger indexes than the size of your dataset

python dask

asked Nov 19 at 13:20

LetsPlayYahtzee

1,89222036

asked Nov 19 at 13:20

LetsPlayYahtzee

1,89222036

asked Nov 19 at 13:20

LetsPlayYahtzee

1,89222036

asked Nov 19 at 13:20

LetsPlayYahtzee

1,89222036

asked Nov 19 at 13:20

LetsPlayYahtzee

1,89222036

add a comment |

active

oldest

votes

Your Answer

StackExchange.ifUsing("editor", function () {
StackExchange.using("externalEditor", function () {
StackExchange.using("snippets", function () {
StackExchange.snippets.init();
});
});
}, "code-snippets");

StackExchange.ready(function() {
var channelOptions = {
tags: "".split(" "),
id: "1"
};
initTagRenderer("".split(" "), "".split(" "), channelOptions);

StackExchange.using("externalEditor", function() {
// Have to fire editor after snippets, if snippets enabled
if (StackExchange.settings.snippets.snippetsEnabled) {
StackExchange.using("snippets", function() {
createEditor();
});
}
else {
createEditor();
}
});

function createEditor() {
StackExchange.prepareEditor({
heartbeatType: 'answer',
convertImagesToLinks: true,
noModals: true,
showLowRepImageUploadWarning: true,
reputationToPostImages: 10,
bindNavPrevention: true,
postfix: "",
imageUploader: {
brandingHtml: "Powered by u003ca class="icon-imgur-white" href="https://imgur.com/"u003eu003c/au003e",
contentPolicyHtml: "User contributions licensed under u003ca href="https://creativecommons.org/licenses/by-sa/3.0/"u003ecc by-sa 3.0 with attribution requiredu003c/au003e u003ca href="https://stackoverflow.com/legal/content-policy"u003e(content policy)u003c/au003e",
allowUrls: true
},
onDemand: true,
discardSelector: ".discard-answer"
,immediatelyShowMarkdownHelp:true
});

}
});

draft saved

draft discarded

Sign up or log in

StackExchange.ready(function () {
StackExchange.helpers.onClickDraftSave('#login-link');
});

Post as a guest

Name

Required, but never shown

StackExchange.ready(
function () {
StackExchange.openid.initPostLogin('.new-post-login', 'https%3a%2f%2fstackoverflow.com%2fquestions%2f53375530%2fhow-to-produce-a-permutation-of-a-dask-dataframe%23new-answer', 'question_page');
}
);

Post as a guest

Name

Required, but never shown

active

oldest

votes

draft saved

draft discarded

draft saved

draft discarded

Sign up or log in

StackExchange.ready(function () {
StackExchange.helpers.onClickDraftSave('#login-link');
});

Post as a guest

Name

Required, but never shown

Post as a guest

Name

Required, but never shown

Sign up or log in

StackExchange.ready(function () {
StackExchange.helpers.onClickDraftSave('#login-link');
});

Post as a guest

Name

Required, but never shown

Sign up or log in

StackExchange.ready(function () {
StackExchange.helpers.onClickDraftSave('#login-link');
});

Post as a guest

Name

Required, but never shown

Sign up or log in

StackExchange.ready(function () {
StackExchange.helpers.onClickDraftSave('#login-link');
});

Post as a guest

Name

Required, but never shown

Name

Required, but never shown

Name

Required, but never shown

This page is only for reference, If you need detailed information, please check here

搜尋此網誌

Ytukyg