mirror of
https://github.com/johnkerl/miller.git
synced 2026-07-29 02:30:12 +00:00
Regularize articultion of Sphinx code blocks (#591)
This commit is contained in:
parent
a5cc45c174
commit
8f0281e2bf
41 changed files with 865 additions and 890 deletions
|
|
@ -224,38 +224,38 @@ What's a CSV file, really? It's an array of rows, or *records*, each being a lis
|
|||
|
||||
For example, if you have:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
shape,flag,index
|
||||
circle,1,24
|
||||
square,0,36
|
||||
GENRST_CARDIFY
|
||||
shape,flag,index
|
||||
circle,1,24
|
||||
square,0,36
|
||||
GENRST_EOF
|
||||
|
||||
then that's a way of saying:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
shape=circle,flag=1,index=24
|
||||
shape=square,flag=0,index=36
|
||||
GENRST_CARDIFY
|
||||
shape=circle,flag=1,index=24
|
||||
shape=square,flag=0,index=36
|
||||
GENRST_EOF
|
||||
|
||||
Other ways to write the same data:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
CSV PPRINT JSON
|
||||
shape,flag,index shape flag index [
|
||||
circle,1,24 circle 1 24 {
|
||||
square,0,36 square 0 36 "shape": "circle",
|
||||
"flag": 1,
|
||||
"index": 24
|
||||
},
|
||||
DKVP XTAB {
|
||||
shape=circle,flag=1,index=24 shape circle "shape": "square",
|
||||
shape=square,flag=0,index=36 flag 1 "flag": 0,
|
||||
index 24 "index": 36
|
||||
}
|
||||
shape square ]
|
||||
flag 0
|
||||
index 36
|
||||
GENRST_CARDIFY
|
||||
CSV PPRINT JSON
|
||||
shape,flag,index shape flag index [
|
||||
circle,1,24 circle 1 24 {
|
||||
square,0,36 square 0 36 "shape": "circle",
|
||||
"flag": 1,
|
||||
"index": 24
|
||||
},
|
||||
DKVP XTAB {
|
||||
shape=circle,flag=1,index=24 shape circle "shape": "square",
|
||||
shape=square,flag=0,index=36 flag 1 "flag": 0,
|
||||
index 24 "index": 36
|
||||
}
|
||||
shape square ]
|
||||
flag 0
|
||||
index 36
|
||||
GENRST_EOF
|
||||
|
||||
Anything we can do with CSV input data, we can do with any other format input data. And you can read from one format, do any record-processing, and output to the same format as the input, or to a different output format.
|
||||
|
||||
|
|
|
|||
|
|
@ -408,8 +408,8 @@ div.highlight { /* CHANGED */
|
|||
|
||||
div.highlight-none { /* CHANGED */
|
||||
background-color: #eae2cb;
|
||||
border-top: 0px solid #ffffff;
|
||||
border-bottom: 10px solid #ffffff;
|
||||
border-top: 2px solid #ffffff;
|
||||
border-bottom: 8px solid #ffffff;
|
||||
border-left: 0px solid #ffffff;
|
||||
border-right: 0px solid #ffffff;
|
||||
padding-left: 0px solid #ffffff;
|
||||
|
|
|
|||
|
|
@ -87,22 +87,22 @@ In this example I am using version 6.1.0 to 6.2.0; of course that will change fo
|
|||
* Similarly for ``macports``: https://github.com/macports/macports-ports/blob/master/textproc/miller/Portfile.
|
||||
* Social-media updates.
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
git remote add upstream https://github.com/Homebrew/homebrew-core # one-time setup only
|
||||
git fetch upstream
|
||||
git rebase upstream/master
|
||||
git checkout -b miller-6.1.0
|
||||
shasum -a 256 /path/to/mlr-6.1.0.tar.gz
|
||||
edit Formula/miller.rb
|
||||
# Test the URL from the line like
|
||||
# url "https://github.com/johnkerl/miller/releases/download/v6.1.0/mlr-6.1.0.tar.gz"
|
||||
# in a browser for typos
|
||||
# A '@BrewTestBot Test this please' comment within the homebrew-core pull request will restart the homebrew travis build
|
||||
git add Formula/miller.rb
|
||||
git commit -m 'miller 6.1.0'
|
||||
git push -u origin miller-6.1.0
|
||||
(submit the pull request)
|
||||
GENRST_CARDIFY
|
||||
git remote add upstream https://github.com/Homebrew/homebrew-core # one-time setup only
|
||||
git fetch upstream
|
||||
git rebase upstream/master
|
||||
git checkout -b miller-6.1.0
|
||||
shasum -a 256 /path/to/mlr-6.1.0.tar.gz
|
||||
edit Formula/miller.rb
|
||||
# Test the URL from the line like
|
||||
# url "https://github.com/johnkerl/miller/releases/download/v6.1.0/mlr-6.1.0.tar.gz"
|
||||
# in a browser for typos
|
||||
# A '@BrewTestBot Test this please' comment within the homebrew-core pull request will restart the homebrew travis build
|
||||
git add Formula/miller.rb
|
||||
git commit -m 'miller 6.1.0'
|
||||
git push -u origin miller-6.1.0
|
||||
(submit the pull request)
|
||||
GENRST_EOF
|
||||
|
||||
* Afterwork:
|
||||
|
||||
|
|
|
|||
|
|
@ -535,7 +535,7 @@ This, of course, depends highly on what's in your log files. But, as an example,
|
|||
I prefer to pre-filter with ``grep`` and/or ``sed`` to extract the structured text, then hand that to Miller. Example:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
grep 'various sorts' *.log | sed 's/.*} //' | mlr --fs space --repifs --oxtab stats1 -a min,p10,p50,p90,max -f time -g status
|
||||
|
||||
|
|
|
|||
|
|
@ -172,16 +172,15 @@ Parsing log-file output
|
|||
|
||||
This, of course, depends highly on what's in your log files. But, as an example, suppose you have log-file lines such as
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
2015-10-08 08:29:09,445 INFO com.company.path.to.ClassName @ [sometext] various/sorts/of data {& punctuation} hits=1 status=0 time=2.378
|
||||
GENRST_CARDIFY
|
||||
2015-10-08 08:29:09,445 INFO com.company.path.to.ClassName @ [sometext] various/sorts/of data {& punctuation} hits=1 status=0 time=2.378
|
||||
GENRST_EOF
|
||||
|
||||
I prefer to pre-filter with ``grep`` and/or ``sed`` to extract the structured text, then hand that to Miller. Example:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
grep 'various sorts' *.log | sed 's/.*} //' | mlr --fs space --repifs --oxtab stats1 -a min,p10,p50,p90,max -f time -g status
|
||||
GENRST_SHOW_COMMAND
|
||||
grep 'various sorts' *.log | sed 's/.*} //' | mlr --fs space --repifs --oxtab stats1 -a min,p10,p50,p90,max -f time -g status
|
||||
GENRST_EOF
|
||||
|
||||
.. _cookbook-memoization-with-oosvars:
|
||||
|
||||
|
|
@ -194,37 +193,37 @@ GENRST_INCLUDE_ESCAPED(data/fibo-uncached.sh)
|
|||
|
||||
produces output like this:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
i o fcount seconds_delta
|
||||
1 1 1 0
|
||||
2 2 3 0.000039101
|
||||
3 3 5 0.000015974
|
||||
4 5 9 0.000019073
|
||||
5 8 15 0.000026941
|
||||
6 13 25 0.000036955
|
||||
7 21 41 0.000056028
|
||||
8 34 67 0.000086069
|
||||
9 55 109 0.000134945
|
||||
10 89 177 0.000217915
|
||||
11 144 287 0.000355959
|
||||
12 233 465 0.000506163
|
||||
13 377 753 0.000811815
|
||||
14 610 1219 0.001297235
|
||||
15 987 1973 0.001960993
|
||||
16 1597 3193 0.003417969
|
||||
17 2584 5167 0.006215811
|
||||
18 4181 8361 0.008294106
|
||||
19 6765 13529 0.012095928
|
||||
20 10946 21891 0.019592047
|
||||
21 17711 35421 0.031193972
|
||||
22 28657 57313 0.057254076
|
||||
23 46368 92735 0.080307961
|
||||
24 75025 150049 0.129482031
|
||||
25 121393 242785 0.213325977
|
||||
26 196418 392835 0.334423065
|
||||
27 317811 635621 0.605969906
|
||||
28 514229 1028457 0.971235037
|
||||
GENRST_CARDIFY
|
||||
i o fcount seconds_delta
|
||||
1 1 1 0
|
||||
2 2 3 0.000039101
|
||||
3 3 5 0.000015974
|
||||
4 5 9 0.000019073
|
||||
5 8 15 0.000026941
|
||||
6 13 25 0.000036955
|
||||
7 21 41 0.000056028
|
||||
8 34 67 0.000086069
|
||||
9 55 109 0.000134945
|
||||
10 89 177 0.000217915
|
||||
11 144 287 0.000355959
|
||||
12 233 465 0.000506163
|
||||
13 377 753 0.000811815
|
||||
14 610 1219 0.001297235
|
||||
15 987 1973 0.001960993
|
||||
16 1597 3193 0.003417969
|
||||
17 2584 5167 0.006215811
|
||||
18 4181 8361 0.008294106
|
||||
19 6765 13529 0.012095928
|
||||
20 10946 21891 0.019592047
|
||||
21 17711 35421 0.031193972
|
||||
22 28657 57313 0.057254076
|
||||
23 46368 92735 0.080307961
|
||||
24 75025 150049 0.129482031
|
||||
25 121393 242785 0.213325977
|
||||
26 196418 392835 0.334423065
|
||||
27 317811 635621 0.605969906
|
||||
28 514229 1028457 0.971235037
|
||||
GENRST_EOF
|
||||
|
||||
Note that the time it takes to evaluate the function is blowing up exponentially as the input argument increases. Using ``@``-variables, which persist across records, we can cache and reuse the results of previous computations:
|
||||
|
||||
|
|
@ -232,34 +231,34 @@ GENRST_INCLUDE_ESCAPED(data/fibo-cached.sh)
|
|||
|
||||
with output like this:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
i o fcount seconds_delta
|
||||
1 1 1 0
|
||||
2 2 3 0.000053883
|
||||
3 3 3 0.000035048
|
||||
4 5 3 0.000045061
|
||||
5 8 3 0.000014067
|
||||
6 13 3 0.000028849
|
||||
7 21 3 0.000028133
|
||||
8 34 3 0.000027895
|
||||
9 55 3 0.000014067
|
||||
10 89 3 0.000015020
|
||||
11 144 3 0.000012875
|
||||
12 233 3 0.000033140
|
||||
13 377 3 0.000014067
|
||||
14 610 3 0.000012875
|
||||
15 987 3 0.000029087
|
||||
16 1597 3 0.000013828
|
||||
17 2584 3 0.000013113
|
||||
18 4181 3 0.000012875
|
||||
19 6765 3 0.000013113
|
||||
20 10946 3 0.000012875
|
||||
21 17711 3 0.000013113
|
||||
22 28657 3 0.000013113
|
||||
23 46368 3 0.000015974
|
||||
24 75025 3 0.000012875
|
||||
25 121393 3 0.000013113
|
||||
26 196418 3 0.000012875
|
||||
27 317811 3 0.000013113
|
||||
28 514229 3 0.000012875
|
||||
GENRST_CARDIFY
|
||||
i o fcount seconds_delta
|
||||
1 1 1 0
|
||||
2 2 3 0.000053883
|
||||
3 3 3 0.000035048
|
||||
4 5 3 0.000045061
|
||||
5 8 3 0.000014067
|
||||
6 13 3 0.000028849
|
||||
7 21 3 0.000028133
|
||||
8 34 3 0.000027895
|
||||
9 55 3 0.000014067
|
||||
10 89 3 0.000015020
|
||||
11 144 3 0.000012875
|
||||
12 233 3 0.000033140
|
||||
13 377 3 0.000014067
|
||||
14 610 3 0.000012875
|
||||
15 987 3 0.000029087
|
||||
16 1597 3 0.000013828
|
||||
17 2584 3 0.000013113
|
||||
18 4181 3 0.000012875
|
||||
19 6765 3 0.000013113
|
||||
20 10946 3 0.000012875
|
||||
21 17711 3 0.000013113
|
||||
22 28657 3 0.000013113
|
||||
23 46368 3 0.000015974
|
||||
24 75025 3 0.000012875
|
||||
25 121393 3 0.000013113
|
||||
26 196418 3 0.000012875
|
||||
27 317811 3 0.000013113
|
||||
28 514229 3 0.000012875
|
||||
GENRST_EOF
|
||||
|
|
|
|||
|
|
@ -10,9 +10,9 @@ Randomly selecting words from a list
|
|||
Given this `word list <./data/english-words.txt>`_, first take a look to see what the first few lines look like:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ head data/english-words.txt
|
||||
head data/english-words.txt
|
||||
a
|
||||
aa
|
||||
aal
|
||||
|
|
@ -27,9 +27,9 @@ Given this `word list <./data/english-words.txt>`_, first take a look to see wha
|
|||
Then the following will randomly sample ten words with four to eight characters in them:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --from data/english-words.txt --nidx filter -S 'n=strlen($1);4<=n&&n<=8' then sample -k 10
|
||||
mlr --from data/english-words.txt --nidx filter -S 'n=strlen($1);4<=n&&n<=8' then sample -k 10
|
||||
thionine
|
||||
birchman
|
||||
mildewy
|
||||
|
|
@ -49,9 +49,9 @@ These are simple *n*-grams as `described here <http://johnkerl.org/randspell/ran
|
|||
The idea is that words from the input file are consumed, then taken apart and pasted back together in ways which imitate the letter-to-letter transitions found in the word list -- giving us automatically generated words in the same vein as *bromance* and *spork*:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --nidx --from ./ngrams/gsl-2000.txt put -q -f ./ngrams/ngfuncs.mlr -f ./ngrams/ng5.mlr
|
||||
mlr --nidx --from ./ngrams/gsl-2000.txt put -q -f ./ngrams/ngfuncs.mlr -f ./ngrams/ng5.mlr
|
||||
beard
|
||||
plastinguish
|
||||
politicially
|
||||
|
|
|
|||
|
|
@ -6,37 +6,35 @@ Randomly selecting words from a list
|
|||
|
||||
Given this `word list <./data/english-words.txt>`_, first take a look to see what the first few lines look like:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ head data/english-words.txt
|
||||
a
|
||||
aa
|
||||
aal
|
||||
aalii
|
||||
aam
|
||||
aardvark
|
||||
aardwolf
|
||||
aba
|
||||
abac
|
||||
abaca
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
head data/english-words.txt
|
||||
a
|
||||
aa
|
||||
aal
|
||||
aalii
|
||||
aam
|
||||
aardvark
|
||||
aardwolf
|
||||
aba
|
||||
abac
|
||||
abaca
|
||||
GENRST_EOF
|
||||
|
||||
Then the following will randomly sample ten words with four to eight characters in them:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ mlr --from data/english-words.txt --nidx filter -S 'n=strlen($1);4<=n&&n<=8' then sample -k 10
|
||||
thionine
|
||||
birchman
|
||||
mildewy
|
||||
avigate
|
||||
addedly
|
||||
abaze
|
||||
askant
|
||||
aiming
|
||||
insulant
|
||||
coinmate
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --from data/english-words.txt --nidx filter -S 'n=strlen($1);4<=n&&n<=8' then sample -k 10
|
||||
thionine
|
||||
birchman
|
||||
mildewy
|
||||
avigate
|
||||
addedly
|
||||
abaze
|
||||
askant
|
||||
aiming
|
||||
insulant
|
||||
coinmate
|
||||
GENRST_EOF
|
||||
|
||||
Randomly generating jabberwocky words
|
||||
----------------------------------------------------------------
|
||||
|
|
@ -45,26 +43,25 @@ These are simple *n*-grams as `described here <http://johnkerl.org/randspell/ran
|
|||
|
||||
The idea is that words from the input file are consumed, then taken apart and pasted back together in ways which imitate the letter-to-letter transitions found in the word list -- giving us automatically generated words in the same vein as *bromance* and *spork*:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ mlr --nidx --from ./ngrams/gsl-2000.txt put -q -f ./ngrams/ngfuncs.mlr -f ./ngrams/ng5.mlr
|
||||
beard
|
||||
plastinguish
|
||||
politicially
|
||||
noise
|
||||
loan
|
||||
country
|
||||
controductionary
|
||||
suppery
|
||||
lose
|
||||
lessors
|
||||
dollar
|
||||
judge
|
||||
rottendence
|
||||
lessenger
|
||||
diffendant
|
||||
suggestional
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --nidx --from ./ngrams/gsl-2000.txt put -q -f ./ngrams/ngfuncs.mlr -f ./ngrams/ng5.mlr
|
||||
beard
|
||||
plastinguish
|
||||
politicially
|
||||
noise
|
||||
loan
|
||||
country
|
||||
controductionary
|
||||
suppery
|
||||
lose
|
||||
lessors
|
||||
dollar
|
||||
judge
|
||||
rottendence
|
||||
lessenger
|
||||
diffendant
|
||||
suggestional
|
||||
GENRST_EOF
|
||||
|
||||
Program timing
|
||||
----------------------------------------------------------------
|
||||
|
|
@ -147,13 +144,13 @@ GENRST_EOF
|
|||
|
||||
But using a very small font size (as small as my Mac will let me go), and by choosing the coordinates to zoom in on a particular part of the complex plane, we can get a nice little picture:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
#!/bin/bash
|
||||
# Get the number of rows and columns from the terminal window dimensions
|
||||
iheight=$(stty size | mlr --nidx --fs space cut -f 1)
|
||||
iwidth=$(stty size | mlr --nidx --fs space cut -f 2)
|
||||
echo "rcorn=-1.755350,icorn=+0.014230,side=0.000020,maxits=10000,iheight=$iheight,iwidth=$iwidth" \
|
||||
| mlr put -f programs/mand.mlr
|
||||
GENRST_CARDIFY
|
||||
#!/bin/bash
|
||||
# Get the number of rows and columns from the terminal window dimensions
|
||||
iheight=$(stty size | mlr --nidx --fs space cut -f 1)
|
||||
iwidth=$(stty size | mlr --nidx --fs space cut -f 2)
|
||||
echo "rcorn=-1.755350,icorn=+0.014230,side=0.000020,maxits=10000,iheight=$iheight,iwidth=$iwidth" \
|
||||
| mlr put -f programs/mand.mlr
|
||||
GENRST_EOF
|
||||
|
||||
.. image:: pix/mand.png
|
||||
|
|
|
|||
|
|
@ -88,12 +88,17 @@ Headerless CSV with duplicate field values
|
|||
|
||||
Miller is (by central design) a mapping from name to value, rather than integer position to value as in most tools in the Unix toolkit such as ``sort``, ``cut``, ``awk``, etc. So given input ``Yea=1,Yea=2`` on the same input line, first ``Yea=1`` is stored, then updated with ``Yea=2``. This is in the input-parser and the value ``Yea=1`` is unavailable to any further processing. The following example line comes from a headerless CSV file and includes 5 times the string (value) ``'NA'``:
|
||||
|
||||
``ag '0.9' nas.csv |head -1
|
||||
2:-349801.10097848,4537221.43295653,2,1,NA,NA,NA,NA,NA``
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
ag '0.9' nas.csv | head -1
|
||||
2:-349801.10097848,4537221.43295653,2,1,NA,NA,NA,NA,NA
|
||||
|
||||
The repeated ``'NA'`` strings (values) in the same line will be treated as fields (columns) with same name, thus only one is kept in the output.
|
||||
|
||||
This can be worked around by telling ``mlr`` that there is no header row by using ``--implicit-csv-header`` or changing the input format by using ``nidx`` like so::
|
||||
This can be worked around by telling ``mlr`` that there is no header row by using ``--implicit-csv-header`` or changing the input format by using ``nidx`` like so:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
ag '0.9' nas.csv | mlr --n2c --fs "," label xsn,ysn,x,y,t,a,e29,e31,e32 then head
|
||||
|
||||
|
|
|
|||
|
|
@ -43,14 +43,18 @@ Headerless CSV with duplicate field values
|
|||
|
||||
Miller is (by central design) a mapping from name to value, rather than integer position to value as in most tools in the Unix toolkit such as ``sort``, ``cut``, ``awk``, etc. So given input ``Yea=1,Yea=2`` on the same input line, first ``Yea=1`` is stored, then updated with ``Yea=2``. This is in the input-parser and the value ``Yea=1`` is unavailable to any further processing. The following example line comes from a headerless CSV file and includes 5 times the string (value) ``'NA'``:
|
||||
|
||||
``ag '0.9' nas.csv |head -1
|
||||
2:-349801.10097848,4537221.43295653,2,1,NA,NA,NA,NA,NA``
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
ag '0.9' nas.csv | head -1
|
||||
2:-349801.10097848,4537221.43295653,2,1,NA,NA,NA,NA,NA
|
||||
GENRST_EOF
|
||||
|
||||
The repeated ``'NA'`` strings (values) in the same line will be treated as fields (columns) with same name, thus only one is kept in the output.
|
||||
|
||||
This can be worked around by telling ``mlr`` that there is no header row by using ``--implicit-csv-header`` or changing the input format by using ``nidx`` like so::
|
||||
This can be worked around by telling ``mlr`` that there is no header row by using ``--implicit-csv-header`` or changing the input format by using ``nidx`` like so:
|
||||
|
||||
ag '0.9' nas.csv | mlr --n2c --fs "," label xsn,ysn,x,y,t,a,e29,e31,e32 then head
|
||||
GENRST_CARDIFY
|
||||
ag '0.9' nas.csv | mlr --n2c --fs "," label xsn,ysn,x,y,t,a,e29,e31,e32 then head
|
||||
GENRST_EOF
|
||||
|
||||
Regularizing ragged CSV
|
||||
----------------------------------------------------------------
|
||||
|
|
|
|||
|
|
@ -10,12 +10,12 @@ How to use .mlrrc
|
|||
Suppose you always use CSV files. Then instead of always having to type ``--csv`` as in
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --csv cut -x -f extra mydata.csv
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --csv sort -n id mydata.csv
|
||||
|
||||
|
|
@ -23,17 +23,17 @@ and so on, you can instead put the following into your ``$HOME/.mlrrc``:
|
|||
|
||||
.. code-block:: none
|
||||
|
||||
--csv
|
||||
--csv
|
||||
|
||||
Then you can just type things like
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr cut -x -f extra mydata.csv
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr sort -n id mydata.csv
|
||||
|
||||
|
|
@ -54,12 +54,10 @@ What you can put in your .mlrrc
|
|||
|
||||
* Empty lines are ignored -- including lines which are empty after comments are removed.
|
||||
|
||||
Here is an example ``.mlrrc file``:
|
||||
Here is an example ``.mlrrc`` file:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
# These are my preferred default settings for Miller
|
||||
|
||||
# Input and output formats are CSV by default (unless otherwise specified
|
||||
# on the mlr command line):
|
||||
csv
|
||||
|
|
|
|||
|
|
@ -6,33 +6,29 @@ How to use .mlrrc
|
|||
|
||||
Suppose you always use CSV files. Then instead of always having to type ``--csv`` as in
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --csv cut -x -f extra mydata.csv
|
||||
GENRST_EOF
|
||||
|
||||
mlr --csv cut -x -f extra mydata.csv
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --csv sort -n id mydata.csv
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --csv sort -n id mydata.csv
|
||||
GENRST_EOF
|
||||
|
||||
and so on, you can instead put the following into your ``$HOME/.mlrrc``:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
GENRST_CARDIFY
|
||||
--csv
|
||||
GENRST_EOF
|
||||
|
||||
Then you can just type things like
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr cut -x -f extra mydata.csv
|
||||
GENRST_EOF
|
||||
|
||||
mlr cut -x -f extra mydata.csv
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr sort -n id mydata.csv
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr sort -n id mydata.csv
|
||||
GENRST_EOF
|
||||
|
||||
and the ``--csv`` part will automatically be understood. (If you do want to process, say, a JSON file then ``mlr --json ...`` at the command line will override the default from your ``.mlrrc``.)
|
||||
|
||||
|
|
@ -51,7 +47,7 @@ What you can put in your .mlrrc
|
|||
|
||||
* Empty lines are ignored -- including lines which are empty after comments are removed.
|
||||
|
||||
Here is an example ``.mlrrc file``:
|
||||
Here is an example ``.mlrrc`` file:
|
||||
|
||||
GENRST_INCLUDE_ESCAPED(sample_mlrrc)
|
||||
|
||||
|
|
|
|||
|
|
@ -135,11 +135,11 @@ Such data are easy to generate, e.g. in Ruby with
|
|||
|
||||
.. code-block:: none
|
||||
|
||||
puts "host=#{hostname},seconds=#{t2-t1},message=#{msg}"
|
||||
puts "host=#{hostname},seconds=#{t2-t1},message=#{msg}"
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
puts mymap.collect{|k,v| "#{k}=#{v}"}.join(',')
|
||||
puts mymap.collect{|k,v| "#{k}=#{v}"}.join(',')
|
||||
|
||||
or ``print`` statements in various languages, e.g.
|
||||
|
||||
|
|
|
|||
|
|
@ -63,33 +63,33 @@ GENRST_EOF
|
|||
|
||||
Such data are easy to generate, e.g. in Ruby with
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_CARDIFY
|
||||
puts "host=#{hostname},seconds=#{t2-t1},message=#{msg}"
|
||||
GENRST_EOF
|
||||
|
||||
puts "host=#{hostname},seconds=#{t2-t1},message=#{msg}"
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
puts mymap.collect{|k,v| "#{k}=#{v}"}.join(',')
|
||||
GENRST_CARDIFY
|
||||
puts mymap.collect{|k,v| "#{k}=#{v}"}.join(',')
|
||||
GENRST_EOF
|
||||
|
||||
or ``print`` statements in various languages, e.g.
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_CARDIFY
|
||||
echo "type=3,user=$USER,date=$date\n";
|
||||
GENRST_EOF
|
||||
|
||||
echo "type=3,user=$USER,date=$date\n";
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
logger.log("type=3,user=$USER,date=$date\n");
|
||||
GENRST_CARDIFY
|
||||
logger.log("type=3,user=$USER,date=$date\n");
|
||||
GENRST_EOF
|
||||
|
||||
Fields lacking an IPS will have positional index (starting at 1) used as the key, as in NIDX format. For example, ``dish=7,egg=8,flint`` is parsed as ``"dish" => "7", "egg" => "8", "3" => "flint"`` and ``dish,egg,flint`` is parsed as ``"1" => "dish", "2" => "egg", "3" => "flint"``.
|
||||
|
||||
As discussed in :doc:`record-heterogeneity`, Miller handles changes of field names within the same data stream. But using DKVP format this is particularly natural. One of my favorite use-cases for Miller is in application/server logs, where I log all sorts of lines such as
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
resource=/path/to/file,loadsec=0.45,ok=true
|
||||
record_count=100, resource=/path/to/file
|
||||
resource=/some/other/path,loadsec=0.97,ok=false
|
||||
GENRST_CARDIFY
|
||||
resource=/path/to/file,loadsec=0.45,ok=true
|
||||
record_count=100, resource=/path/to/file
|
||||
resource=/some/other/path,loadsec=0.97,ok=false
|
||||
GENRST_EOF
|
||||
|
||||
etc. and I just log them as needed. Then later, I can use ``grep``, ``mlr --opprint group-like``, etc.
|
||||
to analyze my logs.
|
||||
|
|
|
|||
|
|
@ -29,6 +29,15 @@ def main
|
|||
cmd_lines = read_until_genrst_eof(input_handle)
|
||||
run_command(cmd_lines, output_handle)
|
||||
|
||||
elsif content_line =~ /^GENRST_CARDIFY$/
|
||||
lines = read_until_genrst_eof(input_handle)
|
||||
write_card([], lines, output_handle)
|
||||
|
||||
elsif content_line =~ /^GENRST_CARDIFY_HIGHLIGHT_ONE$/
|
||||
lines = read_until_genrst_eof(input_handle)
|
||||
line1 = lines.shift
|
||||
write_card([line1], lines, output_handle)
|
||||
|
||||
elsif content_line =~ /^GENRST_RUN_COMMAND_TOLERATING_ERROR$/
|
||||
cmd_lines = read_until_genrst_eof(input_handle)
|
||||
run_command_tolerating_error(cmd_lines, output_handle)
|
||||
|
|
|
|||
|
|
@ -14,38 +14,38 @@ Until then, please see the following sections for how to get Miller 6.*
|
|||
`Homebrew <https://brew.sh/>`_ installation support for OSX is available via
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
brew update && brew install miller
|
||||
|
||||
...and also via `MacPorts <https://www.macports.org/>`_:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
sudo port selfupdate && sudo port install miller
|
||||
|
||||
You may already have the ``mlr`` executable available in your platform's package manager on NetBSD, Debian Linux, Ubuntu Xenial and upward, Arch Linux, or perhaps other distributions. For example, on various Linux distributions you might do one of the following:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
sudo apt-get install miller
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
sudo apt install miller
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
sudo yum install miller
|
||||
|
||||
On Windows, Miller is available via `Chocolatey <https://chocolatey.org/>`_:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
choco install miller
|
||||
|
||||
|
|
|
|||
|
|
@ -10,41 +10,35 @@ Until then, please see the following sections for how to get Miller 6.*
|
|||
|
||||
`Homebrew <https://brew.sh/>`_ installation support for OSX is available via
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
brew update && brew install miller
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
brew update && brew install miller
|
||||
GENRST_EOF
|
||||
|
||||
...and also via `MacPorts <https://www.macports.org/>`_:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
sudo port selfupdate && sudo port install miller
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
sudo port selfupdate && sudo port install miller
|
||||
GENRST_EOF
|
||||
|
||||
You may already have the ``mlr`` executable available in your platform's package manager on NetBSD, Debian Linux, Ubuntu Xenial and upward, Arch Linux, or perhaps other distributions. For example, on various Linux distributions you might do one of the following:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
sudo apt-get install miller
|
||||
GENRST_EOF
|
||||
|
||||
sudo apt-get install miller
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
sudo apt install miller
|
||||
GENRST_EOF
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
sudo apt install miller
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
sudo yum install miller
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
sudo yum install miller
|
||||
GENRST_EOF
|
||||
|
||||
On Windows, Miller is available via `Chocolatey <https://chocolatey.org/>`_:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
choco install miller
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
choco install miller
|
||||
GENRST_EOF
|
||||
|
||||
Prebuilt executables via GitHub per release
|
||||
----------------------------------------------------------------
|
||||
|
|
|
|||
|
|
@ -127,7 +127,9 @@ Miller 6 has getoptish command-line parsing (https://github.com/johnkerl/miller/
|
|||
Improved error messages for DSL parsing
|
||||
----------------------------------------------------------------
|
||||
|
||||
For ``mlr put`` and ``mlr filter``, parse-error messages now include location information::
|
||||
For ``mlr put`` and ``mlr filter``, parse-error messages now include location information:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
mlr: cannot parse DSL expression.
|
||||
Parse error on token ">" at line 63 columnn 7.
|
||||
|
|
|
|||
|
|
@ -115,10 +115,12 @@ Miller 6 has getoptish command-line parsing (https://github.com/johnkerl/miller/
|
|||
Improved error messages for DSL parsing
|
||||
----------------------------------------------------------------
|
||||
|
||||
For ``mlr put`` and ``mlr filter``, parse-error messages now include location information::
|
||||
For ``mlr put`` and ``mlr filter``, parse-error messages now include location information:
|
||||
|
||||
mlr: cannot parse DSL expression.
|
||||
Parse error on token ">" at line 63 columnn 7.
|
||||
GENRST_CARDIFY
|
||||
mlr: cannot parse DSL expression.
|
||||
Parse error on token ">" at line 63 columnn 7.
|
||||
GENRST_EOF
|
||||
|
||||
Developer-specific aspects
|
||||
----------------------------------------------------------------
|
||||
|
|
|
|||
|
|
@ -7,70 +7,70 @@ Quick examples
|
|||
Column select:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --csv cut -f hostname,uptime mydata.csv
|
||||
|
||||
Add new columns as function of other columns:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --nidx put '$sum = $7 < 0.0 ? 3.5 : $7 + 2.1*$8' *.dat
|
||||
|
||||
Row filter:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --csv filter '$status != "down" && $upsec >= 10000' *.csv
|
||||
|
||||
Apply column labels and pretty-print:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
grep -v '^#' /etc/group | mlr --ifs : --nidx --opprint label group,pass,gid,member then sort -f group
|
||||
|
||||
Join multiple data sources on key columns:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr join -j account_id -f accounts.dat then group-by account_name balances.dat
|
||||
|
||||
Mulltiple formats including JSON:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --json put '$attr = sub($attr, "([0-9]+)_([0-9]+)_.*", "\1:\2")' data/*.json
|
||||
|
||||
Aggregate per-column statistics:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr stats1 -a min,mean,max,p10,p50,p90 -f flag,u,v data/*
|
||||
|
||||
Linear regression:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr stats2 -a linreg-pca -f u,v -g shape data/*
|
||||
|
||||
Aggregate custom per-column statistics:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put -q '@sum[$a][$b] += $x; end {emit @sum, "a", "b"}' data/*
|
||||
|
||||
Iterate over data using DSL expressions:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --from estimates.tbl put '
|
||||
for (k,v in $*) {
|
||||
|
|
@ -84,34 +84,34 @@ Iterate over data using DSL expressions:
|
|||
Run DSL expressions from a script file:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --from infile.dat put -f analyze.mlr
|
||||
|
||||
Split/reduce output to multiple filenames:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --from infile.dat put 'tee > "./taps/data-".$a."-".$b, $*'
|
||||
|
||||
Compressed I/O:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --from infile.dat put 'tee | "gzip > ./taps/data-".$a."-".$b.".gz", $*'
|
||||
|
||||
Interoperate with other data-processing tools using standard pipes:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --from infile.dat put -q '@v=$*; dump | "jq .[]"'
|
||||
|
||||
Tap/trace:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --from infile.dat put '(NR % 1000 == 0) { print > stderr, "Checkpoint ".NR}'
|
||||
|
|
|
|||
|
|
@ -3,112 +3,97 @@ Quick examples
|
|||
|
||||
Column select:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --csv cut -f hostname,uptime mydata.csv
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr --csv cut -f hostname,uptime mydata.csv
|
||||
GENRST_EOF
|
||||
|
||||
Add new columns as function of other columns:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --nidx put '$sum = $7 < 0.0 ? 3.5 : $7 + 2.1*$8' *.dat
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --nidx put '$sum = $7 < 0.0 ? 3.5 : $7 + 2.1*$8' *.dat
|
||||
GENRST_EOF
|
||||
|
||||
Row filter:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --csv filter '$status != "down" && $upsec >= 10000' *.csv
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --csv filter '$status != "down" && $upsec >= 10000' *.csv
|
||||
GENRST_EOF
|
||||
|
||||
Apply column labels and pretty-print:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
grep -v '^#' /etc/group | mlr --ifs : --nidx --opprint label group,pass,gid,member then sort -f group
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
grep -v '^#' /etc/group | mlr --ifs : --nidx --opprint label group,pass,gid,member then sort -f group
|
||||
GENRST_EOF
|
||||
|
||||
Join multiple data sources on key columns:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr join -j account_id -f accounts.dat then group-by account_name balances.dat
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr join -j account_id -f accounts.dat then group-by account_name balances.dat
|
||||
GENRST_EOF
|
||||
|
||||
Mulltiple formats including JSON:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --json put '$attr = sub($attr, "([0-9]+)_([0-9]+)_.*", "\1:\2")' data/*.json
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --json put '$attr = sub($attr, "([0-9]+)_([0-9]+)_.*", "\1:\2")' data/*.json
|
||||
GENRST_EOF
|
||||
|
||||
Aggregate per-column statistics:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr stats1 -a min,mean,max,p10,p50,p90 -f flag,u,v data/*
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr stats1 -a min,mean,max,p10,p50,p90 -f flag,u,v data/*
|
||||
GENRST_EOF
|
||||
|
||||
Linear regression:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr stats2 -a linreg-pca -f u,v -g shape data/*
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr stats2 -a linreg-pca -f u,v -g shape data/*
|
||||
GENRST_EOF
|
||||
|
||||
Aggregate custom per-column statistics:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr put -q '@sum[$a][$b] += $x; end {emit @sum, "a", "b"}' data/*
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr put -q '@sum[$a][$b] += $x; end {emit @sum, "a", "b"}' data/*
|
||||
GENRST_EOF
|
||||
|
||||
Iterate over data using DSL expressions:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --from estimates.tbl put '
|
||||
for (k,v in $*) {
|
||||
if (is_numeric(v) && k =~ "^[t-z].*$") {
|
||||
$sum += v; $count += 1
|
||||
}
|
||||
}
|
||||
$mean = $sum / $count # no assignment if count unset
|
||||
'
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --from estimates.tbl put '
|
||||
for (k,v in $*) {
|
||||
if (is_numeric(v) && k =~ "^[t-z].*$") {
|
||||
$sum += v; $count += 1
|
||||
}
|
||||
}
|
||||
$mean = $sum / $count # no assignment if count unset
|
||||
'
|
||||
GENRST_EOF
|
||||
|
||||
Run DSL expressions from a script file:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --from infile.dat put -f analyze.mlr
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --from infile.dat put -f analyze.mlr
|
||||
GENRST_EOF
|
||||
|
||||
Split/reduce output to multiple filenames:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --from infile.dat put 'tee > "./taps/data-".$a."-".$b, $*'
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --from infile.dat put 'tee > "./taps/data-".$a."-".$b, $*'
|
||||
GENRST_EOF
|
||||
|
||||
Compressed I/O:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --from infile.dat put 'tee | "gzip > ./taps/data-".$a."-".$b.".gz", $*'
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --from infile.dat put 'tee | "gzip > ./taps/data-".$a."-".$b.".gz", $*'
|
||||
GENRST_EOF
|
||||
|
||||
Interoperate with other data-processing tools using standard pipes:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --from infile.dat put -q '@v=$*; dump | "jq .[]"'
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --from infile.dat put -q '@v=$*; dump | "jq .[]"'
|
||||
GENRST_EOF
|
||||
|
||||
Tap/trace:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr --from infile.dat put '(NR % 1000 == 0) { print > stderr, "Checkpoint ".NR}'
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --from infile.dat put '(NR % 1000 == 0) { print > stderr, "Checkpoint ".NR}'
|
||||
GENRST_EOF
|
||||
|
|
|
|||
|
|
@ -76,10 +76,12 @@ If-statements
|
|||
These are again reminiscent of ``awk``. Pattern-action blocks are a special case of ``if`` with no ``elif`` or ``else`` blocks, no ``if`` keyword, and parentheses optional around the boolean expression:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put 'NR == 4 {$foo = "bar"}'
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put 'if (NR == 4) {$foo = "bar"}'
|
||||
|
||||
|
|
|
|||
|
|
@ -45,13 +45,13 @@ If-statements
|
|||
|
||||
These are again reminiscent of ``awk``. Pattern-action blocks are a special case of ``if`` with no ``elif`` or ``else`` blocks, no ``if`` keyword, and parentheses optional around the boolean expression:
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr put 'NR == 4 {$foo = "bar"}'
|
||||
GENRST_EOF
|
||||
|
||||
mlr put 'NR == 4 {$foo = "bar"}'
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
mlr put 'if (NR == 4) {$foo = "bar"}'
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr put 'if (NR == 4) {$foo = "bar"}'
|
||||
GENRST_EOF
|
||||
|
||||
Compound statements use ``elif`` (rather than ``elsif`` or ``else if``):
|
||||
|
||||
|
|
|
|||
|
|
@ -6,26 +6,26 @@ Operator precedence
|
|||
|
||||
Operators are listed in order of decreasing precedence, highest first.
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
Operators Associativity
|
||||
--------- -------------
|
||||
() left to right
|
||||
** right to left
|
||||
! ~ unary+ unary- & right to left
|
||||
binary* / // % left to right
|
||||
binary+ binary- . left to right
|
||||
<< >> left to right
|
||||
& left to right
|
||||
^ left to right
|
||||
| left to right
|
||||
< <= > >= left to right
|
||||
== != =~ !=~ left to right
|
||||
&& left to right
|
||||
^^ left to right
|
||||
|| left to right
|
||||
? : right to left
|
||||
= N/A for Miller (there is no $a=$b=$c)
|
||||
GENRST_CARDIFY
|
||||
Operators Associativity
|
||||
--------- -------------
|
||||
() left to right
|
||||
** right to left
|
||||
! ~ unary+ unary- & right to left
|
||||
binary* / // % left to right
|
||||
binary+ binary- . left to right
|
||||
<< >> left to right
|
||||
& left to right
|
||||
^ left to right
|
||||
| left to right
|
||||
< <= > >= left to right
|
||||
== != =~ !=~ left to right
|
||||
&& left to right
|
||||
^^ left to right
|
||||
|| left to right
|
||||
? : right to left
|
||||
= N/A for Miller (there is no $a=$b=$c)
|
||||
GENRST_EOF
|
||||
|
||||
Operator and function semantics
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
|
|
|||
|
|
@ -202,19 +202,19 @@ Semicolons are required between statements even if those statements are on separ
|
|||
Bodies for all compound statements must be enclosed in **curly braces**, even if the body is a single statement:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put 'if ($x == 1) $y = 2' # Syntax error
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put 'if ($x == 1) { $y = 2 }' # This is OK
|
||||
|
||||
Bodies for compound statements may be empty:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put 'if ($x == 1) { }' # This no-op is syntactically acceptable
|
||||
|
||||
|
|
|
|||
|
|
@ -85,20 +85,17 @@ GENRST_INCLUDE_AND_RUN_ESCAPED(data/trailing-commas.sh)
|
|||
|
||||
Bodies for all compound statements must be enclosed in **curly braces**, even if the body is a single statement:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr put 'if ($x == 1) $y = 2' # Syntax error
|
||||
GENRST_EOF
|
||||
|
||||
mlr put 'if ($x == 1) $y = 2' # Syntax error
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr put 'if ($x == 1) { $y = 2 }' # This is OK
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr put 'if ($x == 1) { $y = 2 }' # This is OK
|
||||
GENRST_EOF
|
||||
|
||||
Bodies for compound statements may be empty:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr put 'if ($x == 1) { }' # This no-op is syntactically acceptable
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr put 'if ($x == 1) { }' # This no-op is syntactically acceptable
|
||||
GENRST_EOF
|
||||
|
||||
|
|
|
|||
|
|
@ -271,70 +271,70 @@ Local variables can be defined either untyped as in ``x = 1``, or typed as in ``
|
|||
|
||||
The reason for ``num`` is that ``int`` and ``float`` typedecls are very precise:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
float a = 0; # Runtime error since 0 is int not float
|
||||
int b = 1.0; # Runtime error since 1.0 is float not int
|
||||
num c = 0; # OK
|
||||
num d = 1.0; # OK
|
||||
GENRST_CARDIFY
|
||||
float a = 0; # Runtime error since 0 is int not float
|
||||
int b = 1.0; # Runtime error since 1.0 is float not int
|
||||
num c = 0; # OK
|
||||
num d = 1.0; # OK
|
||||
GENRST_EOF
|
||||
|
||||
A suggestion is to use ``num`` for general use when you want numeric content, and use ``int`` when you genuinely want integer-only values, e.g. in loop indices or map keys (since Miller map keys can only be strings or ints).
|
||||
|
||||
The ``var`` type declaration indicates no type restrictions, e.g. ``var x = 1`` has the same type restrictions on ``x`` as ``x = 1``. The difference is in intentional shadowing: if you have ``x = 1`` in outer scope and ``x = 2`` in inner scope (e.g. within a for-loop or an if-statement) then outer-scope ``x`` has value 2 after the second assignment. But if you have ``var x = 2`` in the inner scope, then you are declaring a variable scoped to the inner block.) For example:
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_CARDIFY
|
||||
x = 1;
|
||||
if (NR == 4) {
|
||||
x = 2; # Refers to outer-scope x: value changes from 1 to 2.
|
||||
}
|
||||
print x; # Value of x is now two
|
||||
GENRST_EOF
|
||||
|
||||
x = 1;
|
||||
if (NR == 4) {
|
||||
x = 2; # Refers to outer-scope x: value changes from 1 to 2.
|
||||
}
|
||||
print x; # Value of x is now two
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
x = 1;
|
||||
if (NR == 4) {
|
||||
var x = 2; # Defines a new inner-scope x with value 2
|
||||
}
|
||||
print x; # Value of this x is still 1
|
||||
GENRST_CARDIFY
|
||||
x = 1;
|
||||
if (NR == 4) {
|
||||
var x = 2; # Defines a new inner-scope x with value 2
|
||||
}
|
||||
print x; # Value of this x is still 1
|
||||
GENRST_EOF
|
||||
|
||||
Likewise function arguments can optionally be typed, with type enforced when the function is called:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
func f(map m, int i) {
|
||||
...
|
||||
}
|
||||
$a = f({1:2, 3:4}, 5); # OK
|
||||
$b = f({1:2, 3:4}, "abc"); # Runtime error
|
||||
$c = f({1:2, 3:4}, $x); # Runtime error for records with non-integer field named x
|
||||
if (NR == 4) {
|
||||
var x = 2; # Defines a new inner-scope x with value 2
|
||||
}
|
||||
print x; # Value of this x is still 1
|
||||
GENRST_CARDIFY
|
||||
func f(map m, int i) {
|
||||
...
|
||||
}
|
||||
$a = f({1:2, 3:4}, 5); # OK
|
||||
$b = f({1:2, 3:4}, "abc"); # Runtime error
|
||||
$c = f({1:2, 3:4}, $x); # Runtime error for records with non-integer field named x
|
||||
if (NR == 4) {
|
||||
var x = 2; # Defines a new inner-scope x with value 2
|
||||
}
|
||||
print x; # Value of this x is still 1
|
||||
GENRST_EOF
|
||||
|
||||
Thirdly, function return values can be type-checked at the point of ``return`` using ``:`` and a typedecl after the parameter list:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
func f(map m, int i): bool {
|
||||
...
|
||||
...
|
||||
if (...) {
|
||||
return "false"; # Runtime error if this branch is taken
|
||||
}
|
||||
...
|
||||
...
|
||||
if (...) {
|
||||
return retval; # Runtime error if this function doesn't have an in-scope
|
||||
# boolean-valued variable named retval
|
||||
}
|
||||
...
|
||||
...
|
||||
# In Miller if your functions don't explicitly return a value, they return absent-null.
|
||||
# So it would also be a runtime error on reaching the end of this function without
|
||||
# an explicit return statement.
|
||||
}
|
||||
GENRST_CARDIFY
|
||||
func f(map m, int i): bool {
|
||||
...
|
||||
...
|
||||
if (...) {
|
||||
return "false"; # Runtime error if this branch is taken
|
||||
}
|
||||
...
|
||||
...
|
||||
if (...) {
|
||||
return retval; # Runtime error if this function doesn't have an in-scope
|
||||
# boolean-valued variable named retval
|
||||
}
|
||||
...
|
||||
...
|
||||
# In Miller if your functions don't explicitly return a value, they return absent-null.
|
||||
# So it would also be a runtime error on reaching the end of this function without
|
||||
# an explicit return statement.
|
||||
}
|
||||
GENRST_EOF
|
||||
|
||||
Null data: empty and absent
|
||||
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
|
|
|
|||
|
|
@ -22,18 +22,18 @@ The short of it is that Miller does this transparently for you so you needn't th
|
|||
|
||||
Implementation details of this, for the interested: integer adds and subtracts overflow by at most one bit so it suffices to check sign-changes. Thus, Miller allows you to add and subtract arbitrary 64-bit signed integers, converting only to float precisely when the result is less than -2\ :sup:`63` or greater than 2\ :sup:`63`\ -1. Multiplies, on the other hand, can overflow by a word size and a sign-change technique does not suffice to detect overflow. Instead Miller tests whether the floating-point product exceeds the representable integer range. Now, 64-bit integers have 64-bit precision while IEEE-doubles have only 52-bit mantissas -- so, there are 53 bits including implicit leading one. The following experiment explicitly demonstrates the resolution at this range:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
64-bit integer 64-bit integer Casted to double Back to 64-bit
|
||||
in hex in decimal integer
|
||||
0x7ffffffffffff9ff 9223372036854774271 9223372036854773760.000000 0x7ffffffffffff800
|
||||
0x7ffffffffffffa00 9223372036854774272 9223372036854773760.000000 0x7ffffffffffff800
|
||||
0x7ffffffffffffbff 9223372036854774783 9223372036854774784.000000 0x7ffffffffffffc00
|
||||
0x7ffffffffffffc00 9223372036854774784 9223372036854774784.000000 0x7ffffffffffffc00
|
||||
0x7ffffffffffffdff 9223372036854775295 9223372036854774784.000000 0x7ffffffffffffc00
|
||||
0x7ffffffffffffe00 9223372036854775296 9223372036854775808.000000 0x8000000000000000
|
||||
0x7ffffffffffffffe 9223372036854775806 9223372036854775808.000000 0x8000000000000000
|
||||
0x7fffffffffffffff 9223372036854775807 9223372036854775808.000000 0x8000000000000000
|
||||
GENRST_CARDIFY
|
||||
64-bit integer 64-bit integer Casted to double Back to 64-bit
|
||||
in hex in decimal integer
|
||||
0x7ffffffffffff9ff 9223372036854774271 9223372036854773760.000000 0x7ffffffffffff800
|
||||
0x7ffffffffffffa00 9223372036854774272 9223372036854773760.000000 0x7ffffffffffff800
|
||||
0x7ffffffffffffbff 9223372036854774783 9223372036854774784.000000 0x7ffffffffffffc00
|
||||
0x7ffffffffffffc00 9223372036854774784 9223372036854774784.000000 0x7ffffffffffffc00
|
||||
0x7ffffffffffffdff 9223372036854775295 9223372036854774784.000000 0x7ffffffffffffc00
|
||||
0x7ffffffffffffe00 9223372036854775296 9223372036854775808.000000 0x8000000000000000
|
||||
0x7ffffffffffffffe 9223372036854775806 9223372036854775808.000000 0x8000000000000000
|
||||
0x7fffffffffffffff 9223372036854775807 9223372036854775808.000000 0x8000000000000000
|
||||
GENRST_EOF
|
||||
|
||||
That is, one cannot check an integer product to see if it is precisely greater than 2\ :sup:`63`\ -1 or less than -2\ :sup:`63` using either integer arithmetic (it may have already overflowed) or using double-precision (due to granularity). Instead Miller checks for overflow in 64-bit integer multiplication by seeing whether the absolute value of the double-precision product exceeds the largest representable IEEE double less than 2\ :sup:`63`, which we see from the listing above is 9223372036854774784. (An alternative would be to do all integer multiplies using handcrafted multi-word 128-bit arithmetic. This approach is not taken.)
|
||||
|
||||
|
|
|
|||
|
|
@ -6,15 +6,15 @@ Formats
|
|||
|
||||
Options:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
--dkvp --idkvp --odkvp
|
||||
--nidx --inidx --onidx
|
||||
--csv --icsv --ocsv
|
||||
--csvlite --icsvlite --ocsvlite
|
||||
--pprint --ipprint --opprint --right
|
||||
--xtab --ixtab --oxtab
|
||||
--json --ijson --ojson
|
||||
GENRST_CARDIFY
|
||||
--dkvp --idkvp --odkvp
|
||||
--nidx --inidx --onidx
|
||||
--csv --icsv --ocsv
|
||||
--csvlite --icsvlite --ocsvlite
|
||||
--pprint --ipprint --opprint --right
|
||||
--xtab --ixtab --oxtab
|
||||
--json --ijson --ojson
|
||||
GENRST_EOF
|
||||
|
||||
These are as discussed in :doc:`file-formats`, with the exception of ``--right`` which makes pretty-printed output right-aligned:
|
||||
|
||||
|
|
@ -50,35 +50,35 @@ Compression
|
|||
|
||||
Options:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
--prepipe {command}
|
||||
GENRST_CARDIFY
|
||||
--prepipe {command}
|
||||
GENRST_EOF
|
||||
|
||||
|
||||
The prepipe command is anything which reads from standard input and produces data acceptable to Miller. Nominally this allows you to use whichever decompression utilities you have installed on your system, on a per-file basis. If the command has flags, quote them: e.g. ``mlr --prepipe 'zcat -cf'``. Examples:
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_CARDIFY
|
||||
# These two produce the same output:
|
||||
$ gunzip < myfile1.csv.gz | mlr cut -f hostname,uptime
|
||||
$ mlr --prepipe gunzip cut -f hostname,uptime myfile1.csv.gz
|
||||
# With multiple input files you need --prepipe:
|
||||
$ mlr --prepipe gunzip cut -f hostname,uptime myfile1.csv.gz myfile2.csv.gz
|
||||
$ mlr --prepipe gunzip --idkvp --oxtab cut -f hostname,uptime myfile1.dat.gz myfile2.dat.gz
|
||||
GENRST_EOF
|
||||
|
||||
# These two produce the same output:
|
||||
$ gunzip < myfile1.csv.gz | mlr cut -f hostname,uptime
|
||||
$ mlr --prepipe gunzip cut -f hostname,uptime myfile1.csv.gz
|
||||
# With multiple input files you need --prepipe:
|
||||
$ mlr --prepipe gunzip cut -f hostname,uptime myfile1.csv.gz myfile2.csv.gz
|
||||
$ mlr --prepipe gunzip --idkvp --oxtab cut -f hostname,uptime myfile1.dat.gz myfile2.dat.gz
|
||||
GENRST_CARDIFY
|
||||
# Similar to the above, but with compressed output as well as input:
|
||||
$ gunzip < myfile1.csv.gz | mlr cut -f hostname,uptime | gzip > outfile.csv.gz
|
||||
$ mlr --prepipe gunzip cut -f hostname,uptime myfile1.csv.gz | gzip > outfile.csv.gz
|
||||
$ mlr --prepipe gunzip cut -f hostname,uptime myfile1.csv.gz myfile2.csv.gz | gzip > outfile.csv.gz
|
||||
GENRST_EOF
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
# Similar to the above, but with compressed output as well as input:
|
||||
$ gunzip < myfile1.csv.gz | mlr cut -f hostname,uptime | gzip > outfile.csv.gz
|
||||
$ mlr --prepipe gunzip cut -f hostname,uptime myfile1.csv.gz | gzip > outfile.csv.gz
|
||||
$ mlr --prepipe gunzip cut -f hostname,uptime myfile1.csv.gz myfile2.csv.gz | gzip > outfile.csv.gz
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
# Similar to the above, but with different compression tools for input and output:
|
||||
$ gunzip < myfile1.csv.gz | mlr cut -f hostname,uptime | xz -z > outfile.csv.xz
|
||||
$ xz -cd < myfile1.csv.xz | mlr cut -f hostname,uptime | gzip > outfile.csv.xz
|
||||
$ mlr --prepipe 'xz -cd' cut -f hostname,uptime myfile1.csv.xz myfile2.csv.xz | xz -z > outfile.csv.xz
|
||||
GENRST_CARDIFY
|
||||
# Similar to the above, but with different compression tools for input and output:
|
||||
$ gunzip < myfile1.csv.gz | mlr cut -f hostname,uptime | xz -z > outfile.csv.xz
|
||||
$ xz -cd < myfile1.csv.xz | mlr cut -f hostname,uptime | gzip > outfile.csv.xz
|
||||
$ mlr --prepipe 'xz -cd' cut -f hostname,uptime myfile1.csv.xz myfile2.csv.xz | xz -z > outfile.csv.xz
|
||||
GENRST_EOF
|
||||
|
||||
.. _reference-separators:
|
||||
|
||||
|
|
@ -89,11 +89,11 @@ Miller has record separators ``IRS`` and ``ORS``, field separators ``IFS`` and `
|
|||
|
||||
Options:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
--rs --irs --ors
|
||||
--fs --ifs --ofs --repifs
|
||||
--ps --ips --ops
|
||||
GENRST_CARDIFY
|
||||
--rs --irs --ors
|
||||
--fs --ifs --ofs --repifs
|
||||
--ps --ips --ops
|
||||
GENRST_EOF
|
||||
|
||||
* You can change a separator from input to output via e.g. ``--ifs = --ofs :``. Or, you can specify that the same separator is to be used for input and output via e.g. ``--fs :``.
|
||||
|
||||
|
|
@ -110,9 +110,9 @@ Number formatting
|
|||
|
||||
The command-line option ``--ofmt {format string}`` is the global number format for commands which generate numeric output, e.g. ``stats1``, ``stats2``, ``histogram``, and ``step``, as well as ``mlr put``. Examples:
|
||||
|
||||
.. code-block:: none
|
||||
|
||||
--ofmt %.9le --ofmt %.6lf --ofmt %.0lf
|
||||
GENRST_CARDIFY
|
||||
--ofmt %.9le --ofmt %.6lf --ofmt %.0lf
|
||||
GENRST_EOF
|
||||
|
||||
These are just familiar ``printf`` formats applied to double-precision numbers. Please don't use ``%s`` or ``%d``. Additionally, if you use leading width (e.g. ``%18.12lf``) then the output will contain embedded whitespace, which may not be what you want if you pipe the output to something else, particularly CSV. I use Miller's pretty-print format (``mlr --opprint``) to column-align numerical data.
|
||||
|
||||
|
|
|
|||
|
|
@ -55,28 +55,27 @@ Regex captures of the form ``\0`` through ``\9`` are supported as
|
|||
* Captures have in-function context for ``sub`` and ``gsub``. For example, the first ``\1,\2`` pair belong to the first ``sub`` and the second ``\1,\2`` pair belong to the second ``sub``:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put '$b = sub($a, "(..)_(...)", "\2-\1"); $c = sub($a, "(..)_(.)(..)", ":\1:\2:\3")'
|
||||
|
||||
* Captures endure for the entirety of a ``put`` for the ``=~`` and ``!=~`` operators. For example, here the ``\1,\2`` are set by the ``=~`` operator and are used by both subsequent assignment statements:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put '$a =~ "(..)_(....); $b = "left_\1"; $c = "right_\2"'
|
||||
|
||||
* The captures are not retained across multiple puts. For example, here the ``\1,\2`` won't be expanded from the regex capture:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr put '$a =~ "(..)_(....)' then {... something else ...} then put '$b = "left_\1"; $c = "right_\2"'
|
||||
|
||||
* Captures are ignored in ``filter`` for the ``=~`` and ``!=~`` operators. For example, there is no mechanism provided to refer to the first ``(..)`` as ``\1`` or to the second ``(....)`` as ``\2`` in the following filter statement:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr filter '$a =~ "(..)_(....)'
|
||||
|
||||
|
|
|
|||
|
|
@ -44,31 +44,27 @@ Regex captures of the form ``\0`` through ``\9`` are supported as
|
|||
|
||||
* Captures have in-function context for ``sub`` and ``gsub``. For example, the first ``\1,\2`` pair belong to the first ``sub`` and the second ``\1,\2`` pair belong to the second ``sub``:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr put '$b = sub($a, "(..)_(...)", "\2-\1"); $c = sub($a, "(..)_(.)(..)", ":\1:\2:\3")'
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr put '$b = sub($a, "(..)_(...)", "\2-\1"); $c = sub($a, "(..)_(.)(..)", ":\1:\2:\3")'
|
||||
GENRST_EOF
|
||||
|
||||
* Captures endure for the entirety of a ``put`` for the ``=~`` and ``!=~`` operators. For example, here the ``\1,\2`` are set by the ``=~`` operator and are used by both subsequent assignment statements:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr put '$a =~ "(..)_(....); $b = "left_\1"; $c = "right_\2"'
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr put '$a =~ "(..)_(....); $b = "left_\1"; $c = "right_\2"'
|
||||
GENRST_EOF
|
||||
|
||||
* The captures are not retained across multiple puts. For example, here the ``\1,\2`` won't be expanded from the regex capture:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr put '$a =~ "(..)_(....)' then {... something else ...} then put '$b = "left_\1"; $c = "right_\2"'
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr put '$a =~ "(..)_(....)' then {... something else ...} then put '$b = "left_\1"; $c = "right_\2"'
|
||||
GENRST_EOF
|
||||
|
||||
* Captures are ignored in ``filter`` for the ``=~`` and ``!=~`` operators. For example, there is no mechanism provided to refer to the first ``(..)`` as ``\1`` or to the second ``(....)`` as ``\2`` in the following filter statement:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr filter '$a =~ "(..)_(....)'
|
||||
GENRST_CARDIFY
|
||||
mlr filter '$a =~ "(..)_(....)'
|
||||
GENRST_EOF
|
||||
|
||||
* Up to nine matches are supported: ``\1`` through ``\9``, while ``\0`` is the entire match string; ``\15`` is treated as ``\1`` followed by an unrelated ``5``.
|
||||
|
||||
|
|
|
|||
|
|
@ -7,14 +7,14 @@ Reference: then-chaining
|
|||
In accord with the `Unix philosophy <http://en.wikipedia.org/wiki/Unix_philosophy>`_, you can pipe data into or out of Miller. For example:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr cut --complement -f os_version *.dat | mlr sort -f hostname,uptime
|
||||
|
||||
You can, if you like, instead simply chain commands together using the ``then`` keyword:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr cut --complement -f os_version then sort -f hostname,uptime *.dat
|
||||
|
||||
|
|
|
|||
|
|
@ -3,17 +3,15 @@ Reference: then-chaining
|
|||
|
||||
In accord with the `Unix philosophy <http://en.wikipedia.org/wiki/Unix_philosophy>`_, you can pipe data into or out of Miller. For example:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr cut --complement -f os_version *.dat | mlr sort -f hostname,uptime
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr cut --complement -f os_version *.dat | mlr sort -f hostname,uptime
|
||||
GENRST_EOF
|
||||
|
||||
You can, if you like, instead simply chain commands together using the ``then`` keyword:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mlr cut --complement -f os_version then sort -f hostname,uptime *.dat
|
||||
GENRST_SHOW_COMMAND
|
||||
mlr cut --complement -f os_version then sort -f hostname,uptime *.dat
|
||||
GENRST_EOF
|
||||
|
||||
(You can precede the very first verb with ``then``, if you like, for symmetry.)
|
||||
|
||||
|
|
|
|||
|
|
@ -147,10 +147,14 @@ bootstrap
|
|||
|
||||
The canonical use for bootstrap sampling is to put error bars on statistical quantities, such as mean. For example:
|
||||
|
||||
..
|
||||
hard-coded, not live-code, since random sampling would generate different data on each doc run
|
||||
which would needlessly complicate git diff
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --opprint stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
mlr --opprint stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.497129 1413
|
||||
red 0.492560 4641
|
||||
|
|
@ -162,7 +166,7 @@ The canonical use for bootstrap sampling is to put error bars on statistical qua
|
|||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.500651 1380
|
||||
purple 0.501556 1111
|
||||
|
|
@ -174,7 +178,7 @@ The canonical use for bootstrap sampling is to put error bars on statistical qua
|
|||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.498046 1485
|
||||
blue 0.513576 1417
|
||||
|
|
@ -186,7 +190,7 @@ The canonical use for bootstrap sampling is to put error bars on statistical qua
|
|||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
blue 0.522921 1447
|
||||
red 0.490717 4617
|
||||
|
|
|
|||
|
|
@ -78,53 +78,53 @@ GENRST_EOF
|
|||
|
||||
The canonical use for bootstrap sampling is to put error bars on statistical quantities, such as mean. For example:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
..
|
||||
hard-coded, not live-code, since random sampling would generate different data on each doc run
|
||||
which would needlessly complicate git diff
|
||||
|
||||
$ mlr --opprint stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.497129 1413
|
||||
red 0.492560 4641
|
||||
purple 0.494005 1142
|
||||
green 0.504861 1109
|
||||
blue 0.517717 1470
|
||||
orange 0.490532 303
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --opprint stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.497129 1413
|
||||
red 0.492560 4641
|
||||
purple 0.494005 1142
|
||||
green 0.504861 1109
|
||||
blue 0.517717 1470
|
||||
orange 0.490532 303
|
||||
GENRST_EOF
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.500651 1380
|
||||
purple 0.501556 1111
|
||||
green 0.503272 1068
|
||||
red 0.493895 4702
|
||||
blue 0.512529 1496
|
||||
orange 0.521030 321
|
||||
GENRST_EOF
|
||||
|
||||
$ mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.500651 1380
|
||||
purple 0.501556 1111
|
||||
green 0.503272 1068
|
||||
red 0.493895 4702
|
||||
blue 0.512529 1496
|
||||
orange 0.521030 321
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.498046 1485
|
||||
blue 0.513576 1417
|
||||
red 0.492870 4595
|
||||
orange 0.507697 307
|
||||
green 0.496803 1075
|
||||
purple 0.486337 1199
|
||||
GENRST_EOF
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
yellow 0.498046 1485
|
||||
blue 0.513576 1417
|
||||
red 0.492870 4595
|
||||
orange 0.507697 307
|
||||
green 0.496803 1075
|
||||
purple 0.486337 1199
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
blue 0.522921 1447
|
||||
red 0.490717 4617
|
||||
yellow 0.496450 1419
|
||||
purple 0.496523 1192
|
||||
green 0.507569 1111
|
||||
orange 0.468014 292
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --opprint bootstrap then stats1 -a mean,count -f u -g color data/colored-shapes.dkvp
|
||||
color u_mean u_count
|
||||
blue 0.522921 1447
|
||||
red 0.490717 4617
|
||||
yellow 0.496450 1419
|
||||
purple 0.496523 1192
|
||||
green 0.507569 1111
|
||||
orange 0.468014 292
|
||||
GENRST_EOF
|
||||
|
||||
.. _reference-verbs-cat:
|
||||
|
||||
|
|
|
|||
|
|
@ -9,10 +9,10 @@ The Miller REPL (read-evaluate-print loop) is an interactive counterpart to reco
|
|||
It isn't a source-level debugger which lets you execute one source-code *statement* at a time -- however, it does let you operate on one *record* at a time. Further, it lets you use "immediate expressions", namely, you can interact with the language without having to provide data from an input file.
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr repl
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr repl
|
||||
|
||||
[mlr] 1 + 2
|
||||
3
|
||||
|
||||
|
|
@ -103,34 +103,36 @@ Examples
|
|||
Use the REPL to look at arithmetic:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr repl
|
||||
|
||||
mlr repl
|
||||
|
||||
[mlr] 6/3
|
||||
2
|
||||
|
||||
|
||||
[mlr] 6/5
|
||||
1.2
|
||||
|
||||
|
||||
[mlr] typeof(6/3)
|
||||
int
|
||||
|
||||
|
||||
[mlr] typeof(6/5)
|
||||
float
|
||||
|
||||
Read the first record from a small file:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr repl
|
||||
|
||||
mlr repl
|
||||
|
||||
[mlr] :open foo.dat
|
||||
|
||||
|
||||
[mlr] :read
|
||||
|
||||
|
||||
[mlr] :context
|
||||
FILENAME="foo.dat",FILENUM=1,NR=1,FNR=1
|
||||
|
||||
|
||||
[mlr] $*
|
||||
{
|
||||
"a": "eks",
|
||||
|
|
@ -139,18 +141,19 @@ Read the first record from a small file:
|
|||
"x": 0.38139939387114097,
|
||||
"y": 0.13418874328430463
|
||||
}
|
||||
|
||||
|
||||
[mlr] $z = $x + $i
|
||||
|
||||
|
||||
[mlr] :write
|
||||
a=eks,b=wye,i=4,x=0.38139939387114097,y=0.13418874328430463,z=4.381399393871141
|
||||
|
||||
Skip until deep into a larger file, then inspect a record:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr repl --csv
|
||||
|
||||
mlr repl --csv
|
||||
|
||||
[mlr] :open data/colored-shapes.csv
|
||||
[mlr] :skip until NR == 10000
|
||||
[mlr] :r
|
||||
|
|
|
|||
|
|
@ -5,13 +5,12 @@ The Miller REPL (read-evaluate-print loop) is an interactive counterpart to reco
|
|||
|
||||
It isn't a source-level debugger which lets you execute one source-code *statement* at a time -- however, it does let you operate on one *record* at a time. Further, it lets you use "immediate expressions", namely, you can interact with the language without having to provide data from an input file.
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr repl
|
||||
|
||||
$ mlr repl
|
||||
|
||||
[mlr] 1 + 2
|
||||
3
|
||||
[mlr] 1 + 2
|
||||
3
|
||||
GENRST_EOF
|
||||
|
||||
Using Miller without the REPL
|
||||
----------------------------------------------------------------
|
||||
|
|
@ -80,69 +79,69 @@ Examples
|
|||
|
||||
Use the REPL to look at arithmetic:
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr repl
|
||||
|
||||
$ mlr repl
|
||||
[mlr] 6/3
|
||||
2
|
||||
|
||||
[mlr] 6/3
|
||||
2
|
||||
[mlr] 6/5
|
||||
1.2
|
||||
|
||||
[mlr] 6/5
|
||||
1.2
|
||||
[mlr] typeof(6/3)
|
||||
int
|
||||
|
||||
[mlr] typeof(6/3)
|
||||
int
|
||||
|
||||
[mlr] typeof(6/5)
|
||||
float
|
||||
[mlr] typeof(6/5)
|
||||
float
|
||||
GENRST_EOF
|
||||
|
||||
Read the first record from a small file:
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr repl
|
||||
|
||||
$ mlr repl
|
||||
[mlr] :open foo.dat
|
||||
|
||||
[mlr] :open foo.dat
|
||||
[mlr] :read
|
||||
|
||||
[mlr] :read
|
||||
[mlr] :context
|
||||
FILENAME="foo.dat",FILENUM=1,NR=1,FNR=1
|
||||
|
||||
[mlr] :context
|
||||
FILENAME="foo.dat",FILENUM=1,NR=1,FNR=1
|
||||
[mlr] $*
|
||||
{
|
||||
"a": "eks",
|
||||
"b": "wye",
|
||||
"i": 4,
|
||||
"x": 0.38139939387114097,
|
||||
"y": 0.13418874328430463
|
||||
}
|
||||
|
||||
[mlr] $*
|
||||
{
|
||||
"a": "eks",
|
||||
"b": "wye",
|
||||
"i": 4,
|
||||
"x": 0.38139939387114097,
|
||||
"y": 0.13418874328430463
|
||||
}
|
||||
[mlr] $z = $x + $i
|
||||
|
||||
[mlr] $z = $x + $i
|
||||
|
||||
[mlr] :write
|
||||
a=eks,b=wye,i=4,x=0.38139939387114097,y=0.13418874328430463,z=4.381399393871141
|
||||
[mlr] :write
|
||||
a=eks,b=wye,i=4,x=0.38139939387114097,y=0.13418874328430463,z=4.381399393871141
|
||||
GENRST_EOF
|
||||
|
||||
Skip until deep into a larger file, then inspect a record:
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr repl --csv
|
||||
|
||||
$ mlr repl --csv
|
||||
|
||||
[mlr] :open data/colored-shapes.csv
|
||||
[mlr] :skip until NR == 10000
|
||||
[mlr] :r
|
||||
[mlr] $*
|
||||
{
|
||||
"color": "yellow",
|
||||
"shape": "circle",
|
||||
"flag": 1,
|
||||
"i": 99284,
|
||||
"u": 0.6530503199545348,
|
||||
"v": 0.23908588907834516,
|
||||
"w": 0.4799125551304738,
|
||||
"x": 6.379888206335166
|
||||
}
|
||||
[mlr] :open data/colored-shapes.csv
|
||||
[mlr] :skip until NR == 10000
|
||||
[mlr] :r
|
||||
[mlr] $*
|
||||
{
|
||||
"color": "yellow",
|
||||
"shape": "circle",
|
||||
"flag": 1,
|
||||
"i": 99284,
|
||||
"u": 0.6530503199545348,
|
||||
"v": 0.23908588907834516,
|
||||
"w": 0.4799125551304738,
|
||||
"x": 6.379888206335166
|
||||
}
|
||||
GENRST_EOF
|
||||
|
||||
History-editing
|
||||
----------------------------------------------------------------
|
||||
|
|
|
|||
|
|
@ -1,5 +1,3 @@
|
|||
# These are my preferred default settings for Miller
|
||||
|
||||
# Input and output formats are CSV by default (unless otherwise specified
|
||||
# on the mlr command line):
|
||||
csv
|
||||
|
|
|
|||
|
|
@ -22,24 +22,22 @@ Note that running a subprocess on every record takes a non-trivial amount of tim
|
|||
..
|
||||
hard-coded, not live-code, since %N doesn't exist on all platforms
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --opprint put '$t=system("date +%s.%N")' then step -a delta -f t data/small
|
||||
a b i x y t t_delta
|
||||
pan pan 1 0.3467901443380824 0.7268028627434533 1568774318.513903817 0
|
||||
eks pan 2 0.7586799647899636 0.5221511083334797 1568774318.514722876 0.000819
|
||||
wye wye 3 0.20460330576630303 0.33831852551664776 1568774318.515618046 0.000895
|
||||
eks wye 4 0.38139939387114097 0.13418874328430463 1568774318.516547441 0.000929
|
||||
wye pan 5 0.5732889198020006 0.8636244699032729 1568774318.517518828 0.000971
|
||||
GENRST_EOF
|
||||
|
||||
$ mlr --opprint put '$t=system("date +%s.%N")' then step -a delta -f t data/small
|
||||
a b i x y t t_delta
|
||||
pan pan 1 0.3467901443380824 0.7268028627434533 1568774318.513903817 0
|
||||
eks pan 2 0.7586799647899636 0.5221511083334797 1568774318.514722876 0.000819
|
||||
wye wye 3 0.20460330576630303 0.33831852551664776 1568774318.515618046 0.000895
|
||||
eks wye 4 0.38139939387114097 0.13418874328430463 1568774318.516547441 0.000929
|
||||
wye pan 5 0.5732889198020006 0.8636244699032729 1568774318.517518828 0.000971
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ mlr --opprint put '$t=systime()' then step -a delta -f t data/small
|
||||
a b i x y t t_delta
|
||||
pan pan 1 0.3467901443380824 0.7268028627434533 1568774318.518699 0
|
||||
eks pan 2 0.7586799647899636 0.5221511083334797 1568774318.518717 0.000018
|
||||
wye wye 3 0.20460330576630303 0.33831852551664776 1568774318.518723 0.000006
|
||||
eks wye 4 0.38139939387114097 0.13418874328430463 1568774318.518727 0.000004
|
||||
wye pan 5 0.5732889198020006 0.8636244699032729 1568774318.518730 0.000003
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --opprint put '$t=systime()' then step -a delta -f t data/small
|
||||
a b i x y t t_delta
|
||||
pan pan 1 0.3467901443380824 0.7268028627434533 1568774318.518699 0
|
||||
eks pan 2 0.7586799647899636 0.5221511083334797 1568774318.518717 0.000018
|
||||
wye wye 3 0.20460330576630303 0.33831852551664776 1568774318.518723 0.000006
|
||||
eks wye 4 0.38139939387114097 0.13418874328430463 1568774318.518727 0.000004
|
||||
wye pan 5 0.5732889198020006 0.8636244699032729 1568774318.518730 0.000003
|
||||
GENRST_EOF
|
||||
|
|
|
|||
|
|
@ -14,9 +14,9 @@ I like to produce SQL-query output with header-column and tab delimiter: this is
|
|||
For example, using default output formatting in ``mysql`` we get formatting like Miller's ``--opprint --barred``:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mysql --database=mydb -e 'show columns in mytable'
|
||||
mysql --database=mydb -e 'show columns in mytable'
|
||||
+------------------+--------------+------+-----+---------+-------+
|
||||
| Field | Type | Null | Key | Default | Extra |
|
||||
+------------------+--------------+------+-----+---------+-------+
|
||||
|
|
@ -30,9 +30,9 @@ For example, using default output formatting in ``mysql`` we get formatting like
|
|||
Using ``mysql``'s ``-B`` we get TSV output:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mysql --database=mydb -B -e 'show columns in mytable' | mlr --itsvlite --opprint cat
|
||||
mysql --database=mydb -B -e 'show columns in mytable' | mlr --itsvlite --opprint cat
|
||||
Field Type Null Key Default Extra
|
||||
id bigint(20) NO MUL NULL -
|
||||
category varchar(256) NO - NULL -
|
||||
|
|
@ -43,9 +43,9 @@ Using ``mysql``'s ``-B`` we get TSV output:
|
|||
Since Miller handles TSV output, we can do as much or as little processing as we want in the SQL query, then send the rest on to Miller. This includes outputting as JSON, doing further selects/joins in Miller, doing stats, etc. etc.:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mysql --database=mydb -B -e 'show columns in mytable' | mlr --itsvlite --ojson --jlistwrap --jvstack cat
|
||||
mysql --database=mydb -B -e 'show columns in mytable' | mlr --itsvlite --ojson --jlistwrap --jvstack cat
|
||||
[
|
||||
{
|
||||
"Field": "id",
|
||||
|
|
@ -90,11 +90,14 @@ Since Miller handles TSV output, we can do as much or as little processing as we
|
|||
]
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mysql --database=mydb -B -e 'select * from mytable' > query.tsv
|
||||
mysql --database=mydb -B -e 'select * from mytable' > query.tsv
|
||||
|
||||
$ mlr --from query.tsv --t2p stats1 -a count -f id -g category,assigned_to
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
mlr --from query.tsv --t2p stats1 -a count -f id -g category,assigned_to
|
||||
category assigned_to id_count
|
||||
special 10000978 207
|
||||
special 10003924 385
|
||||
|
|
@ -125,13 +128,13 @@ Create and load SQL table:
|
|||
y DOUBLE
|
||||
);
|
||||
Query OK, 0 rows affected (0.01 sec)
|
||||
|
||||
|
||||
bash$ mlr --onidx --fs comma cat data/medium > medium.nidx
|
||||
|
||||
|
||||
mysql> LOAD DATA LOCAL INFILE 'medium.nidx' REPLACE INTO TABLE abixy FIELDS TERMINATED BY ',' ;
|
||||
Query OK, 10000 rows affected (0.07 sec)
|
||||
Records: 10000 Deleted: 0 Skipped: 0 Warnings: 0
|
||||
|
||||
|
||||
mysql> SELECT COUNT(*) AS count FROM abixy;
|
||||
+-------+
|
||||
| count |
|
||||
|
|
@ -139,7 +142,7 @@ Create and load SQL table:
|
|||
| 10000 |
|
||||
+-------+
|
||||
1 row in set (0.00 sec)
|
||||
|
||||
|
||||
mysql> SELECT * FROM abixy LIMIT 10;
|
||||
+------+------+------+---------------------+---------------------+
|
||||
| a | b | i | x | y |
|
||||
|
|
@ -159,7 +162,6 @@ Create and load SQL table:
|
|||
Aggregate counts within SQL:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mysql> SELECT a, b, COUNT(*) AS count FROM abixy GROUP BY a, b ORDER BY COUNT DESC;
|
||||
+------+------+-------+
|
||||
|
|
@ -196,9 +198,9 @@ Aggregate counts within SQL:
|
|||
Aggregate counts within Miller:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mlr --opprint uniq -c -g a,b then sort -nr count data/medium
|
||||
mlr --opprint uniq -c -g a,b then sort -nr count data/medium
|
||||
a b count
|
||||
zee wye 455
|
||||
pan eks 429
|
||||
|
|
@ -219,9 +221,9 @@ Aggregate counts within Miller:
|
|||
Pipe SQL output to aggregate counts within Miller:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
:emphasize-lines: 1-1
|
||||
|
||||
$ mysql -D miller -B -e 'select * from abixy' | mlr --itsv --opprint uniq -c -g a,b then sort -nr count
|
||||
mysql -D miller -B -e 'select * from abixy' | mlr --itsv --opprint uniq -c -g a,b then sort -nr count
|
||||
a b count
|
||||
zee wye 455
|
||||
pan eks 429
|
||||
|
|
|
|||
|
|
@ -10,96 +10,94 @@ I like to produce SQL-query output with header-column and tab delimiter: this is
|
|||
|
||||
For example, using default output formatting in ``mysql`` we get formatting like Miller's ``--opprint --barred``:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ mysql --database=mydb -e 'show columns in mytable'
|
||||
+------------------+--------------+------+-----+---------+-------+
|
||||
| Field | Type | Null | Key | Default | Extra |
|
||||
+------------------+--------------+------+-----+---------+-------+
|
||||
| id | bigint(20) | NO | MUL | NULL | |
|
||||
| category | varchar(256) | NO | | NULL | |
|
||||
| is_permanent | tinyint(1) | NO | | NULL | |
|
||||
| assigned_to | bigint(20) | YES | | NULL | |
|
||||
| last_update_time | int(11) | YES | | NULL | |
|
||||
+------------------+--------------+------+-----+---------+-------+
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mysql --database=mydb -e 'show columns in mytable'
|
||||
+------------------+--------------+------+-----+---------+-------+
|
||||
| Field | Type | Null | Key | Default | Extra |
|
||||
+------------------+--------------+------+-----+---------+-------+
|
||||
| id | bigint(20) | NO | MUL | NULL | |
|
||||
| category | varchar(256) | NO | | NULL | |
|
||||
| is_permanent | tinyint(1) | NO | | NULL | |
|
||||
| assigned_to | bigint(20) | YES | | NULL | |
|
||||
| last_update_time | int(11) | YES | | NULL | |
|
||||
+------------------+--------------+------+-----+---------+-------+
|
||||
GENRST_EOF
|
||||
|
||||
Using ``mysql``'s ``-B`` we get TSV output:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ mysql --database=mydb -B -e 'show columns in mytable' | mlr --itsvlite --opprint cat
|
||||
Field Type Null Key Default Extra
|
||||
id bigint(20) NO MUL NULL -
|
||||
category varchar(256) NO - NULL -
|
||||
is_permanent tinyint(1) NO - NULL -
|
||||
assigned_to bigint(20) YES - NULL -
|
||||
last_update_time int(11) YES - NULL -
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mysql --database=mydb -B -e 'show columns in mytable' | mlr --itsvlite --opprint cat
|
||||
Field Type Null Key Default Extra
|
||||
id bigint(20) NO MUL NULL -
|
||||
category varchar(256) NO - NULL -
|
||||
is_permanent tinyint(1) NO - NULL -
|
||||
assigned_to bigint(20) YES - NULL -
|
||||
last_update_time int(11) YES - NULL -
|
||||
GENRST_EOF
|
||||
|
||||
Since Miller handles TSV output, we can do as much or as little processing as we want in the SQL query, then send the rest on to Miller. This includes outputting as JSON, doing further selects/joins in Miller, doing stats, etc. etc.:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mysql --database=mydb -B -e 'show columns in mytable' | mlr --itsvlite --ojson --jlistwrap --jvstack cat
|
||||
[
|
||||
{
|
||||
"Field": "id",
|
||||
"Type": "bigint(20)",
|
||||
"Null": "NO",
|
||||
"Key": "MUL",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
},
|
||||
{
|
||||
"Field": "category",
|
||||
"Type": "varchar(256)",
|
||||
"Null": "NO",
|
||||
"Key": "",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
},
|
||||
{
|
||||
"Field": "is_permanent",
|
||||
"Type": "tinyint(1)",
|
||||
"Null": "NO",
|
||||
"Key": "",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
},
|
||||
{
|
||||
"Field": "assigned_to",
|
||||
"Type": "bigint(20)",
|
||||
"Null": "YES",
|
||||
"Key": "",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
},
|
||||
{
|
||||
"Field": "last_update_time",
|
||||
"Type": "int(11)",
|
||||
"Null": "YES",
|
||||
"Key": "",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
}
|
||||
]
|
||||
GENRST_EOF
|
||||
|
||||
$ mysql --database=mydb -B -e 'show columns in mytable' | mlr --itsvlite --ojson --jlistwrap --jvstack cat
|
||||
[
|
||||
{
|
||||
"Field": "id",
|
||||
"Type": "bigint(20)",
|
||||
"Null": "NO",
|
||||
"Key": "MUL",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
},
|
||||
{
|
||||
"Field": "category",
|
||||
"Type": "varchar(256)",
|
||||
"Null": "NO",
|
||||
"Key": "",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
},
|
||||
{
|
||||
"Field": "is_permanent",
|
||||
"Type": "tinyint(1)",
|
||||
"Null": "NO",
|
||||
"Key": "",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
},
|
||||
{
|
||||
"Field": "assigned_to",
|
||||
"Type": "bigint(20)",
|
||||
"Null": "YES",
|
||||
"Key": "",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
},
|
||||
{
|
||||
"Field": "last_update_time",
|
||||
"Type": "int(11)",
|
||||
"Null": "YES",
|
||||
"Key": "",
|
||||
"Default": "NULL",
|
||||
"Extra": ""
|
||||
}
|
||||
]
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mysql --database=mydb -B -e 'select * from mytable' > query.tsv
|
||||
GENRST_EOF
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ mysql --database=mydb -B -e 'select * from mytable' > query.tsv
|
||||
|
||||
$ mlr --from query.tsv --t2p stats1 -a count -f id -g category,assigned_to
|
||||
category assigned_to id_count
|
||||
special 10000978 207
|
||||
special 10003924 385
|
||||
special 10009872 168
|
||||
standard 10000978 524
|
||||
standard 10003924 392
|
||||
standard 10009872 108
|
||||
...
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --from query.tsv --t2p stats1 -a count -f id -g category,assigned_to
|
||||
category assigned_to id_count
|
||||
special 10000978 207
|
||||
special 10003924 385
|
||||
special 10009872 168
|
||||
standard 10000978 524
|
||||
standard 10003924 392
|
||||
standard 10009872 108
|
||||
...
|
||||
GENRST_EOF
|
||||
|
||||
Again, all the examples in the CSV section apply here -- just change the input-format flags.
|
||||
|
||||
|
|
@ -112,136 +110,133 @@ One use of NIDX (value-only, no keys) format is for loading up SQL tables.
|
|||
|
||||
Create and load SQL table:
|
||||
|
||||
.. code-block:: none
|
||||
GENRST_CARDIFY
|
||||
mysql> CREATE TABLE abixy(
|
||||
a VARCHAR(32),
|
||||
b VARCHAR(32),
|
||||
i BIGINT(10),
|
||||
x DOUBLE,
|
||||
y DOUBLE
|
||||
);
|
||||
Query OK, 0 rows affected (0.01 sec)
|
||||
|
||||
mysql> CREATE TABLE abixy(
|
||||
a VARCHAR(32),
|
||||
b VARCHAR(32),
|
||||
i BIGINT(10),
|
||||
x DOUBLE,
|
||||
y DOUBLE
|
||||
);
|
||||
Query OK, 0 rows affected (0.01 sec)
|
||||
bash$ mlr --onidx --fs comma cat data/medium > medium.nidx
|
||||
|
||||
bash$ mlr --onidx --fs comma cat data/medium > medium.nidx
|
||||
mysql> LOAD DATA LOCAL INFILE 'medium.nidx' REPLACE INTO TABLE abixy FIELDS TERMINATED BY ',' ;
|
||||
Query OK, 10000 rows affected (0.07 sec)
|
||||
Records: 10000 Deleted: 0 Skipped: 0 Warnings: 0
|
||||
|
||||
mysql> LOAD DATA LOCAL INFILE 'medium.nidx' REPLACE INTO TABLE abixy FIELDS TERMINATED BY ',' ;
|
||||
Query OK, 10000 rows affected (0.07 sec)
|
||||
Records: 10000 Deleted: 0 Skipped: 0 Warnings: 0
|
||||
mysql> SELECT COUNT(*) AS count FROM abixy;
|
||||
+-------+
|
||||
| count |
|
||||
+-------+
|
||||
| 10000 |
|
||||
+-------+
|
||||
1 row in set (0.00 sec)
|
||||
|
||||
mysql> SELECT COUNT(*) AS count FROM abixy;
|
||||
+-------+
|
||||
| count |
|
||||
+-------+
|
||||
| 10000 |
|
||||
+-------+
|
||||
1 row in set (0.00 sec)
|
||||
|
||||
mysql> SELECT * FROM abixy LIMIT 10;
|
||||
+------+------+------+---------------------+---------------------+
|
||||
| a | b | i | x | y |
|
||||
+------+------+------+---------------------+---------------------+
|
||||
| pan | pan | 1 | 0.3467901443380824 | 0.7268028627434533 |
|
||||
| eks | pan | 2 | 0.7586799647899636 | 0.5221511083334797 |
|
||||
| wye | wye | 3 | 0.20460330576630303 | 0.33831852551664776 |
|
||||
| eks | wye | 4 | 0.38139939387114097 | 0.13418874328430463 |
|
||||
| wye | pan | 5 | 0.5732889198020006 | 0.8636244699032729 |
|
||||
| zee | pan | 6 | 0.5271261600918548 | 0.49322128674835697 |
|
||||
| eks | zee | 7 | 0.6117840605678454 | 0.1878849191181694 |
|
||||
| zee | wye | 8 | 0.5985540091064224 | 0.976181385699006 |
|
||||
| hat | wye | 9 | 0.03144187646093577 | 0.7495507603507059 |
|
||||
| pan | wye | 10 | 0.5026260055412137 | 0.9526183602969864 |
|
||||
+------+------+------+---------------------+---------------------+
|
||||
mysql> SELECT * FROM abixy LIMIT 10;
|
||||
+------+------+------+---------------------+---------------------+
|
||||
| a | b | i | x | y |
|
||||
+------+------+------+---------------------+---------------------+
|
||||
| pan | pan | 1 | 0.3467901443380824 | 0.7268028627434533 |
|
||||
| eks | pan | 2 | 0.7586799647899636 | 0.5221511083334797 |
|
||||
| wye | wye | 3 | 0.20460330576630303 | 0.33831852551664776 |
|
||||
| eks | wye | 4 | 0.38139939387114097 | 0.13418874328430463 |
|
||||
| wye | pan | 5 | 0.5732889198020006 | 0.8636244699032729 |
|
||||
| zee | pan | 6 | 0.5271261600918548 | 0.49322128674835697 |
|
||||
| eks | zee | 7 | 0.6117840605678454 | 0.1878849191181694 |
|
||||
| zee | wye | 8 | 0.5985540091064224 | 0.976181385699006 |
|
||||
| hat | wye | 9 | 0.03144187646093577 | 0.7495507603507059 |
|
||||
| pan | wye | 10 | 0.5026260055412137 | 0.9526183602969864 |
|
||||
+------+------+------+---------------------+---------------------+
|
||||
GENRST_EOF
|
||||
|
||||
Aggregate counts within SQL:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
mysql> SELECT a, b, COUNT(*) AS count FROM abixy GROUP BY a, b ORDER BY COUNT DESC;
|
||||
+------+------+-------+
|
||||
| a | b | count |
|
||||
+------+------+-------+
|
||||
| zee | wye | 455 |
|
||||
| pan | eks | 429 |
|
||||
| pan | pan | 427 |
|
||||
| wye | hat | 426 |
|
||||
| hat | wye | 423 |
|
||||
| pan | hat | 417 |
|
||||
| eks | hat | 417 |
|
||||
| pan | zee | 413 |
|
||||
| eks | eks | 413 |
|
||||
| zee | hat | 409 |
|
||||
| eks | wye | 407 |
|
||||
| zee | zee | 403 |
|
||||
| pan | wye | 395 |
|
||||
| wye | pan | 392 |
|
||||
| zee | eks | 391 |
|
||||
| zee | pan | 389 |
|
||||
| hat | eks | 389 |
|
||||
| wye | eks | 386 |
|
||||
| wye | zee | 385 |
|
||||
| hat | zee | 385 |
|
||||
| hat | hat | 381 |
|
||||
| wye | wye | 377 |
|
||||
| eks | pan | 371 |
|
||||
| hat | pan | 363 |
|
||||
| eks | zee | 357 |
|
||||
+------+------+-------+
|
||||
25 rows in set (0.01 sec)
|
||||
GENRST_CARDIFY
|
||||
mysql> SELECT a, b, COUNT(*) AS count FROM abixy GROUP BY a, b ORDER BY COUNT DESC;
|
||||
+------+------+-------+
|
||||
| a | b | count |
|
||||
+------+------+-------+
|
||||
| zee | wye | 455 |
|
||||
| pan | eks | 429 |
|
||||
| pan | pan | 427 |
|
||||
| wye | hat | 426 |
|
||||
| hat | wye | 423 |
|
||||
| pan | hat | 417 |
|
||||
| eks | hat | 417 |
|
||||
| pan | zee | 413 |
|
||||
| eks | eks | 413 |
|
||||
| zee | hat | 409 |
|
||||
| eks | wye | 407 |
|
||||
| zee | zee | 403 |
|
||||
| pan | wye | 395 |
|
||||
| wye | pan | 392 |
|
||||
| zee | eks | 391 |
|
||||
| zee | pan | 389 |
|
||||
| hat | eks | 389 |
|
||||
| wye | eks | 386 |
|
||||
| wye | zee | 385 |
|
||||
| hat | zee | 385 |
|
||||
| hat | hat | 381 |
|
||||
| wye | wye | 377 |
|
||||
| eks | pan | 371 |
|
||||
| hat | pan | 363 |
|
||||
| eks | zee | 357 |
|
||||
+------+------+-------+
|
||||
25 rows in set (0.01 sec)
|
||||
GENRST_EOF
|
||||
|
||||
Aggregate counts within Miller:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ mlr --opprint uniq -c -g a,b then sort -nr count data/medium
|
||||
a b count
|
||||
zee wye 455
|
||||
pan eks 429
|
||||
pan pan 427
|
||||
wye hat 426
|
||||
hat wye 423
|
||||
pan hat 417
|
||||
eks hat 417
|
||||
eks eks 413
|
||||
pan zee 413
|
||||
zee hat 409
|
||||
eks wye 407
|
||||
zee zee 403
|
||||
pan wye 395
|
||||
hat pan 363
|
||||
eks zee 357
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mlr --opprint uniq -c -g a,b then sort -nr count data/medium
|
||||
a b count
|
||||
zee wye 455
|
||||
pan eks 429
|
||||
pan pan 427
|
||||
wye hat 426
|
||||
hat wye 423
|
||||
pan hat 417
|
||||
eks hat 417
|
||||
eks eks 413
|
||||
pan zee 413
|
||||
zee hat 409
|
||||
eks wye 407
|
||||
zee zee 403
|
||||
pan wye 395
|
||||
hat pan 363
|
||||
eks zee 357
|
||||
GENRST_EOF
|
||||
|
||||
Pipe SQL output to aggregate counts within Miller:
|
||||
|
||||
.. code-block:: none
|
||||
:emphasize-lines: 1,1
|
||||
|
||||
$ mysql -D miller -B -e 'select * from abixy' | mlr --itsv --opprint uniq -c -g a,b then sort -nr count
|
||||
a b count
|
||||
zee wye 455
|
||||
pan eks 429
|
||||
pan pan 427
|
||||
wye hat 426
|
||||
hat wye 423
|
||||
pan hat 417
|
||||
eks hat 417
|
||||
eks eks 413
|
||||
pan zee 413
|
||||
zee hat 409
|
||||
eks wye 407
|
||||
zee zee 403
|
||||
pan wye 395
|
||||
wye pan 392
|
||||
zee eks 391
|
||||
zee pan 389
|
||||
hat eks 389
|
||||
wye eks 386
|
||||
hat zee 385
|
||||
wye zee 385
|
||||
hat hat 381
|
||||
wye wye 377
|
||||
eks pan 371
|
||||
hat pan 363
|
||||
eks zee 357
|
||||
GENRST_CARDIFY_HIGHLIGHT_ONE
|
||||
mysql -D miller -B -e 'select * from abixy' | mlr --itsv --opprint uniq -c -g a,b then sort -nr count
|
||||
a b count
|
||||
zee wye 455
|
||||
pan eks 429
|
||||
pan pan 427
|
||||
wye hat 426
|
||||
hat wye 423
|
||||
pan hat 417
|
||||
eks hat 417
|
||||
eks eks 413
|
||||
pan zee 413
|
||||
zee hat 409
|
||||
eks wye 407
|
||||
zee zee 403
|
||||
pan wye 395
|
||||
wye pan 392
|
||||
zee eks 391
|
||||
zee pan 389
|
||||
hat eks 389
|
||||
wye eks 386
|
||||
hat zee 385
|
||||
wye zee 385
|
||||
hat hat 381
|
||||
wye wye 377
|
||||
eks pan 371
|
||||
hat pan 363
|
||||
eks zee 357
|
||||
GENRST_EOF
|
||||
|
|
|
|||
17
go/todo.txt
17
go/todo.txt
|
|
@ -7,11 +7,6 @@ TOP OF LIST:
|
|||
* survey email followups
|
||||
|
||||
* convert all code-blocks to cardify
|
||||
* most/least-frequent: map -> odict
|
||||
|
||||
----------------------------------------------------------------
|
||||
colorizer:
|
||||
* mrpl output?
|
||||
|
||||
----------------------------------------------------------------
|
||||
doclink etc:
|
||||
|
|
@ -19,16 +14,14 @@ doclink etc:
|
|||
* data/small -> csv throughout. and/or just use example.csv
|
||||
----------------------------------------------------------------
|
||||
|
||||
----------------------------------------------------------------
|
||||
HELP
|
||||
|
||||
! type-arithmetic-info
|
||||
|
||||
----------------------------------------------------------------
|
||||
|
||||
----------------------------------------------------------------
|
||||
DOC6
|
||||
|
||||
! 'S mlr'
|
||||
! sphinx bootstrap comment
|
||||
! file:////Users/kerl/pub_http_internet/miller-releases/miller-head/docs6/_build/html/shapes-of-data.html
|
||||
diagnosing margin-upper bork
|
||||
|
||||
* memory page:
|
||||
o open with "out of memory" -- what next?
|
||||
o streaming vs non-streaming
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue