Showing posts with label regex. Show all posts
Showing posts with label regex. Show all posts

Tuesday, February 2, 2016

Use searched pattern in substitution

Sometimes we want to search for a pattern and use that pattern in replace as well. Like the substitution is an extension of what we already have in the file.

E.g. append all employee ages with 'years'
emp_name:Jane
emp_age:25
emp_name:John
emp_age:22
emp_name:Mary
emp_age:30
emp_name:Natalie
emp_age:29
other data here
Command
:%s/\(emp_age.*\)/\1 years/g
Breaking it up -
emp_name:Jane

%s - apply thru out the file
\( and \) - escape the parentheses 
(emp_age.*) - consider it a group
\1 - escape 1; 1 means the first pattern, in our case it's the only one
\1 years -  use the first pattern and append years to it
g - replace multiple occurrences on the same line
Output
emp_name:Jane
emp_age:25 years
emp_name:John
emp_age:22 years
emp_name:Mary
emp_age:30 years
emp_name:Natalie
emp_age:29 years
other data here

While I was typing this up, I thought let me try using a second pattern as well.
:%s/\(emp_age.*\)\(years\)/\1- \2/g 
If we run this in the above output we get
emp_name:Jane
emp_age:25 - years
emp_name:John
emp_age:22 - years
emp_name:Mary
emp_age:30 - years
emp_name:Natalie
emp_age:29 - years
other data here
Hope this helps.

Thursday, January 12, 2012

Regular Expressions in Informatica

Informatica supports PERL like regex syntax. So any parsing you would do in shell, can be done in Informatica as well.


REG_EXTRACT('Fiction__JK_ROWLING__HARRY_POTTER__MAX07','(\w+)(__\w+__)(\w+)(__.*)',3)



In the code above, __ are the delimiters in the string.

The regext part
\w looks for any alphanumeric character including underscore. Adding + to it makes it look for multiple alphanum characters.
() are to group the patterns
. looks for any character and adding & * makes it look for all the occurrences of any characters.
So we have four groups and the output is the 3rd group.

OUTPUT
HARRY_POTTER

Can you figure the other groups? Now can you figure how many groups we need if we just need the number of books in series or just the genre?

Wednesday, May 11, 2011

vi or sed - Delete till end of line from pattern

Here's the regex that will match everything from the pattern until a newline is found.

Say you want to remove everything after "here" from this sample file.

code here    where?
code here   right here!
code here         look every where.
Used regex
s/here.*//
Explanation:
.*  
dot matches any character including spaces, numbers, alphabets but not new line.
asterisk matches any number of occurrences of a character

s/here.*/meow/ will result in

code meow
code meow
code meow