# Tricky regex on HTML

**URL:** <https://forum.kirupa.com/t/tricky-regex-on-html/278887>\
**Category:** Uncategorized\
**Created:** [January 7, 2009, 1:03am UTC](https://forum.kirupa.com/t/tricky-regex-on-html/278887 "2009-01-07T01:03:12Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![NeoDreamer](https://avatars.discourse-cdn.com/v4/letter/n/ccd318/32.png) [@NeoDreamer](https://forum.kirupa.com/u/NeoDreamer)\
**Post date:** [January 7, 2009, 1:03am UTC](https://forum.kirupa.com/t/tricky-regex-on-html/278887/1 "2009-01-07T01:03:12Z")

</div>

I’m trying to use a regex to extract the first author’s name from a thread such as: [http://www.threadless.com/profile/470607/wotto/blog/247841/percentage\_Blowwwwg](http://www.threadless.com/profile/470607/wotto/blog/247841/percentage_Blowwwwg)

The relevant HTML is:

```auto

			&nbsp;
			Aug 01 '07 by 
							
				<a class="lavendar" href="/profile/470607/wotto">wotto</a>
				  		&nbsp;&nbsp;&nbsp;&nbsp;

```

I’ve gotten this far. It matches the date.

```php

preg_match_all('/[
][	]{3}[A-Z]{1}[a-z]{2} [0-9]{2} \'[0-9]{2} by [\r]/', $html, $match);

```

The next step is to add the tab part: []{7} to the end of the regex. But doing so makes the match start failing. It is strange that they used both newline [  
] and return [\r] in the same document. Maybe both Linux and Windows programmers were working on their thread page code. Perhaps Linux has an analogous symbol to [] that I don’t know of?
